Categories
Nevin Manimala Statistics

Evaluation of large language models in root resorption scenarios: an ESE-aligned comparative performance assessment

Odontology. 2026 Aug 10. doi: 10.1007/s10266-026-01531-z. Online ahead of print.

ABSTRACT

This study aims to compare the diagnostic accuracy, appropriateness of treatment planning, and source citation performance of five large language models ChatGPT-4o (Free), ChatGPT-5.1 Plus, Microsoft Copilot, Google Gemini, and DeepSeek-R1 in root resorption scenarios. In December 2025, twelve clinical scenarios were created based on the classification of the European Society of Endodontology and each scenario was presented to all chatbots over four consecutive days. All responses were evaluated using a blinded assessment protocol and a binary scoring system. A total of 720 observations (12 cases × 4 repetitions × 3 criteria per model) were analyzed. The collected data were analyzed using chi-square, Fisher’s exact, and Cochran Q tests. In terms of diagnostic accuracy, Microsoft Copilot (79.2%), ChatGPT-5.1 (77.1%), and ChatGPT-4o (Free) (75%) showed the highest performance. Google Gemini (68.8%) demonstrated a moderate level of accuracy, while DeepSeek (39.6%) showed markedly low performance. All models exhibited high accuracy in treatment plan recommendations, and no statistically significant differences were detected. Regarding citation accuracy, Copilot ranked first with 100% accuracy. Although large language models present potential as supportive decision-making tools in the evaluation of root resorption, diagnostic inconsistencies, limitations in source accuracy, and variability in responses restrict their independent use in clinical applications. Therefore, the outputs generated by these models should be interpreted cautiously within clinical decision-making processes.

PMID:42573919 | DOI:10.1007/s10266-026-01531-z

By Nevin Manimala

Portfolio Website for Nevin Manimala