Categories
Nevin Manimala Statistics

Correctness, Harmfulness, and Diversity of Large Language Models for Colonoscopy Preparation Assistance: Comparative Evaluation Study

JMIR AI. 2026 Aug 4;5:e88581. doi: 10.2196/88581.

ABSTRACT

BACKGROUND: Colorectal cancer is a leading cause of cancer-related deaths in the United States, and colonoscopy remains the gold standard for early detection and prevention. However, many procedures are postponed due to inadequate bowel preparation, a preventable failure often caused by patients’ difficulty in understanding and following written prep instructions. Prior interventions such as reminder apps and instructional videos have improved adherence only modestly, largely because they cannot answer patient-specific questions. Recent advances in large language models (LLMs) raise the possibility of developing conversational assistants that can provide interactive support to patients in procedure preparation.

OBJECTIVE: This study evaluated the correctness, harmfulness, and diversity of synthetic dialogues generated by leading LLMs acting as both simulated AI Coaches and patients for colonoscopy preparation.

METHODS: Five leading LLMs-OpenAI’s o3, GPT-4.1, and GPT-5.1; Meta’s Llama 3.3 70B; and Mistral’s Large-2411-were used to generate 250 patient-AI Coach dialogues per model. Dialogues consisted of 3 to 7 question-answer pairs concerning diet, medications, and other prep-related topics. A multiprompt, multiquestion approach was designed to elicit diverse patient questions, and an error taxonomy was established to assess model capabilities in responding to questions. Human raters, including 3 medical experts, evaluated the generated questions for difficulty and the responses for correctness, error type, and potential harmfulness. Automatic evaluation using an LLM-as-a-judge approach complemented human evaluation. Question diversity was assessed using lexical diversity metrics (Distinct-1 and Distinct-2) and entropy. In addition, we evaluated a safety filtering mechanism in which responses judged incorrect by an automated evaluator were replaced with a deferral message instructing patients to contact their health care provider. Differences in response correctness across models were evaluated using permutation tests conducted at the dialogue level. Interrater agreement among human evaluators was assessed using the Gwet AC1 statistic. The study was conducted between May and September 2025.

RESULTS: Automatic evaluation results closely aligned with human judgments: leading models approached but did not achieve adequate performance. Closed-weight models (GPT-5.1, GPT-4.1, and o3) outperformed open-weight models (Llama and Mistral) on correctness, with the reasoning models (GPT-5.1 and o3) performing best. This turn-level ranking was preserved under the supplementary single-prompt baseline, although dialogue-level rankings differed. All models produced harmful errors, primarily due to omissions or misinterpretations of prep instructions. The multiprompt generation strategy substantially increased the diversity of patient questions compared with a single-prompt baseline. Applying an automated safety filter reduced overall error rates but failed to eliminate harmful responses.

CONCLUSIONS: Although LLMs demonstrate strong potential to support colonoscopy preparation, none are yet reliable enough for unsupervised deployment in patient-facing contexts. Persistent harmful errors and the limited effectiveness of simple filtering mechanisms highlight the need for improved instruction adherence, stronger safety mechanisms, and validation using real patient queries.

PMID:42550998 | DOI:10.2196/88581

By Nevin Manimala

Portfolio Website for Nevin Manimala