Front Digit Health. 2026 Jul 23;8:1877420. doi: 10.3389/fdgth.2026.1877420. eCollection 2026.
ABSTRACT
BACKGROUND/OBJECTIVES: Multimodal large language models are increasingly being explored for medical image interpretation, but their performance in thyroid nodule ultrasound assessment remains uncertain. This study compared ChatGPT-5.4 and DeepSeek-VL2 for C-TIRADS-based risk stratification of thyroid nodules using static ultrasound images.
METHODS: This single-center retrospective diagnostic accuracy study included 206 pathologically confirmed thyroid nodules, including 117 benign and 89 malignant nodules. For each nodule, paired transverse and longitudinal ultrasound images were analyzed by both models using the same standardized English prompt. The models generated structured outputs for C-TIRADS-related ultrasound features and estimated C-TIRADS category. Pathology served as the reference standard for malignancy, and expert-assigned C-TIRADS served as the imaging reference standard. Model-assigned C-TIRADS categories were used as ordinal risk scores. The primary analysis used C-TIRADS ≥4B as the binary positive threshold for predicting pathology-confirmed malignancy, and a sensitivity analysis used C-TIRADS ≥4A.
RESULTS: At the ≥4B threshold, ChatGPT-5.4 achieved higher accuracy, sensitivity, and specificity than DeepSeek-VL2: 70.9%, 61.8%, and 77.8% vs. 57.8%, 50.6%, and 63.2%, respectively. Its non-diagnostic rate was also lower (6.3% vs. 15.5%). In the valid-output ROC analysis, AUCs were 0.903 for expert assessment, 0.790 for ChatGPT-5.4, and 0.757 for DeepSeek-VL2; model-based ROC analyses excluded indeterminate outputs and should be interpreted as per-protocol estimates. At the C-TIRADS category level, ChatGPT-5.4 showed higher exact agreement with expert assessment than DeepSeek-VL2 (45.6% vs. 36.4%) and higher within-one-category agreement (87.6% vs. 82.2%). Feature-level analysis showed that both models performed best on echogenicity but poorly on margin and echogenic foci. At the ≥4A threshold, sensitivity increased but specificity decreased substantially for both models.
CONCLUSIONS: ChatGPT-5.4 showed a more favorable performance profile than DeepSeek-VL2, but differences in sensitivity and valid-output AUC were not statistically significant. Both models remain investigational adjunctive tools and should not be used for unsupervised clinical decision-making or standalone triage.
PMID:42564679 | PMC:PMC13444745 | DOI:10.3389/fdgth.2026.1877420