Categories
Nevin Manimala Statistics

Evaluation of Large Language Model-Generated Recommendations in Glaucoma Surgical Decision-Making

Semin Ophthalmol. 2026 Aug 29:1-7. doi: 10.1080/08820538.2026.2725223. Online ahead of print.

ABSTRACT

OBJECTIVE: To evaluate the utility, rationality, and safety of glaucoma surgery recommendations generated by three prominent large language models (LLMs) – ChatGPT, Microsoft Copilot, and Google Gemini – when applied to real-world clinical scenarios.

METHODS: Retrospective records from a tertiary hospital were converted into standardized scenarios and stratified into “primary” and “complex” glaucoma groups. Each LLM was prompted to suggest a single surgical approach and provide a rationale. A blinded team of glaucoma specialists evaluated the outputs based on six criteria: appropriateness, rationale quality, specificity, adherence to guidelines, feasibility, and safety risk, using a normalized 0-100 scale.

RESULTS: Median overall quality scores across all cases were 80.7 for ChatGPT, 80.7 for Copilot, and 82.7 for Gemini, showing no statistically significant difference in general performance (p = .367). However, case complexity significantly affected performance. For Gemini, appropriateness and rationale quality scores dropped significantly in complex cases and were accompanied by a statistically significant increase in safety risk (p = .009). Although ChatGPT and Copilot demonstrated more stability across groups, their rationale quality was significantly lower in complex scenarios than in primary ones (p = .014 and <0.001, respectively). Pairwise analyses revealed that ChatGPT offered superior rationale quality compared to Copilot, while Gemini exhibited higher specificity.

CONCLUSIONS: LLM-based chatbots can provide acceptable surgical guidance for straightforward primary surgical cases, but their utility is limited in high-risk or complex clinical settings. The observed deficiencies in rationale and increased safety risks in complex cases suggest that LLMs should be integrated as auxiliary decision-support tools under expert supervision rather than used as autonomous decision-makers in glaucoma surgery planning.

PMID:42667272 | DOI:10.1080/08820538.2026.2725223

By Nevin Manimala

Portfolio Website for Nevin Manimala