Comparison of ChatGPT’s Diagnostic and Management Accuracy of Foot and Ankle Bone-Related Pathologies to Orthopaedic Surgeons

J Am Acad Orthop Surg. 2025 Apr 10. doi: 10.5435/JAAOS-D-24-01049. Online ahead of print.

ABSTRACT

INTRODUCTION: The steep rise in utilization of large language model chatbots, such as ChatGPT, has spilled into medicine in recent years. The newest version of ChatGPT, ChatGPT-4, has passed medical licensure examinations and, specifically in orthopaedics, has performed at the level of a postgraduate level three orthopaedic surgery resident on the Orthopaedic In-Service Training Examination question bank sets. The purpose of this study was to evaluate ChatGPT-4’s diagnostic and decision-making capacity in the clinical management of bone-related injuries of the foot and ankle.

METHODS: Eight bone-related foot and ankle orthopaedic cases were presented to ChatGPT-4 and subsequently evaluated by three fellowship-trained foot and ankle orthopaedic surgeons. Cases were scored using criteria on a Likert scale, graded from a total score of 5 (lowest) to 25 (highest) across five criteria. ChatGPT-4 was referred to as “Dr. GPT,” establishing a peer dynamic so that the role of an orthopaedic surgeon was emulated by the chatbot.

RESULTS: The average score across all criteria for each case was 4.53 of 5, noting an overall average sum score of 22.7 of 25 for all cases. The pathology with the highest score was the second metatarsal stress fracture (24.3), whereas the case with the lowest score was hallux rigidus (21.3). Kendall correlation analysis of interrater reliability showed variable correlation among surgeons, without statistical significance.

CONCLUSION: ChatGPT-4 effectively diagnosed and provided appropriate treatment options for simple bone-related foot and ankle cases. Importantly, ChatGPT did not fabricate treatment options (ie, hallucination phenomenon), which has been previously well-documented in the literature, notably receiving its second-highest overall average score in this criterion. ChatGPT struggled to provide comprehensive information beyond standard treatment options. Overall, ChatGPT has the potential to serve as a widely accessible resource for patients and nonorthopaedic clinicians, although limitations may exist in the delivery of comprehensive information.

PMID:40233367 | DOI:10.5435/JAAOS-D-24-01049

By Nevin Manimala