Front Public Health. 2026 Jul 31;14:1915867. doi: 10.3389/fpubh.2026.1915867. eCollection 2026.
ABSTRACT
BACKGROUND: After diagnosis of high-risk human papillomavirus (HPV) infection, patients often seek guidance on cancer risk, colposcopy, follow-up, treatment, partner management, pregnancy, and anxiety. Large language models (LLMs) are increasingly used for medical consultation, but their response quality and safety in this setting require evaluation.
METHODS: In this single-center, expert-rated methodological study, 15 patient-style questions based on common outpatient consultations were submitted to ChatGPT-5.5 Instant and DeepSeek-V3. Each model generated three independent responses per question, yielding 90 artificial intelligence (AI)-generated responses. Five blinded gynecology experts independently evaluated all responses, producing 450 expert-rating records. Accuracy, safety, guideline concordance, completeness, and understandability were rated on a 1-5 Likert scale. Experts also assessed major error and potential harm. The primary outcome was the proportion of responses without major error or potential harm.
RESULTS: ChatGPT-5.5 Instant had higher response-level scores than DeepSeek-V3 for accuracy (mean difference, 0.25; 95% CI, 0.11-0.39; FDR-adjusted p = 0.010), safety (0.27; 0.07-0.47; p = 0.010), completeness (0.22; 0.07-0.37; p = 0.020), and composite score (0.18; 0.04-0.31; p = 0.027). The difference in guideline concordance did not remain significant after multiplicity correction (0.16; -0.01 to 0.33; FDR-adjusted p = 0.057), and understandability was similar. Responses without major error or potential harm occurred in 42/45 (93.3%; 95% CI, 81.7-98.6%) ChatGPT responses and 38/45 (84.4%; 70.5-93.5%) DeepSeek-V3 responses (risk difference, 8.9 percentage points; 95% CI, -13.3 to 31.1; p = 0.344). Exploratory question-level risks were observed in HPV16/18 positivity, normal cytology, partner management, and pregnancy scenarios.
CONCLUSION: Both two models demonstrated generally strong performance across the evaluated quality domains. ChatGPT achieved higher ratings in several quality domains, but response-level binary safety differences were imprecise and not statistically significant. Both models require guideline-based clinical oversight.
PMID:42601968 | PMC:PMC13473150 | DOI:10.3389/fpubh.2026.1915867