Health Informatics J. 2026 Jul-Sep;32(3):14604582261470637. doi: 10.1177/14604582261470637. Epub 2026 Jul 28.
ABSTRACT
ObjectiveThis study evaluates the performance of large language models (LLMs)-ChatGPT-4.0, Gemini 2.0 Pro, o3-mini, Doctor GPT and DeepSeek-V3-in a national orthopaedic proficiency examination and explores their implications for health informatics and medical education. The responses of these models were analysed to assess accuracy rates and differences between models.MethodA total of 100 multiple-choice questions from the 2024 TOTEK examination were administered to each AI model under identical conditions. Correct and incorrect responses were recorded, and differences in performance were evaluated using chi-square testing and frequency analysis. Question categories were also compared to identify domain-specific variations.Resultso3-mini achieved the highest accuracy rate (79%), while Gemini 2.0 showed the lowest (68%); all models exceeded the 60% pass threshold. A statistically significant difference between models was identified in the Surgical Procedures category, in which Gemini 2.0 answered fewer questions correctly (23/36) than the other models (30-32/36) (χ2 = 9.87, df = 4, p = 0.043). No significant differences were observed in the remaining categories (all p > 0.05), and the overall difference in accuracy between models did not reach statistical significance (χ2 = 4.01, df = 4, p = 0.405). Clinical decision-making and visual content-based questions were the most challenging for all models.ConclusionAI models demonstrate generally high accuracy in medical examinations; however, they struggle with interpreting clinical context, recognising atypical medical scenarios and answering questions involving visual content.
PMID:42520276 | DOI:10.1177/14604582261470637