Foot Ankle Surg. 2026 Aug 11:S1268-7731(26)00213-4. doi: 10.1016/j.fas.2026.08.002. Online ahead of print.
ABSTRACT
BACKGROUND: This study compared seven large language models (LLMs) to identify the optimal model for foot and ankle clinical decision support.
METHODS: The LLMs answered 20 multiple-choice questions (MCQs) and 20 open-ended clinical questions. MCQ accuracy was recorded. Three blinded foot and ankle surgeons evaluated the open-ended responses based on accuracy, completeness, and clinical relevance (total score range: 3-21).
RESULTS: GPT-o3 and GPT-5 Thinking achieved the highest MCQ accuracy (95%). For open-ended evaluations, mean total scores differed significantly across the models (p < 0.001). GPT-5 Thinking (19.98 ± 0.66) and GPT-o3 (19.67 ± 0.51) attained the highest scores, significantly outperforming the other five models, with no statistical difference between these top two performers.
CONCLUSION: LLM performance for foot and ankle disorders differed substantially across models. GPT-5 Thinking and GPT-o3 exhibited superior performance under the tested conditions, emphasizing the necessity of targeted model selection for clinical decision-making. However, because LLMs evolve rapidly, these findings should be interpreted as a time-specific snapshot rather than as fixed or generalizable rankings of model performance.
PMID:42595676 | DOI:10.1016/j.fas.2026.08.002