Categories
Nevin Manimala Statistics

Large Language Models in German Continuing Medical Education Assessments: Protocol for a Fully Crossed Experimental Study

JMIR Res Protoc. 2026 Jul 28;15:e91675. doi: 10.2196/91675.

ABSTRACT

BACKGROUND: Continuing medical education (CME) is a legal and ethical obligation for physicians in Germany. The rapid rise of large language models (LLMs) such as ChatGPT, Gemini, Claude, and Grok raises concerns about the integrity of CME assessments, as LLMs can already pass German CME tests.

OBJECTIVE: This study aims to determine whether the choice of document format (searchable PDF, protected PDF, raster PDF, or vector PDF) and LLM influences the ability of LLMs to solve CME test questions at rates exceeding the passing threshold specified for each CME module (typically 70%).

METHODS: In a fully crossed within-subjects repeated-measures design, 18 expired CME articles from 3 major German publishers across 6 specialties will be converted into 3 cheating-impeding PDF formats and processed alongside the original PDF files by 4 current LLMs (GPT-5, Claude Sonnet 4, Grok-4, and Gemini 3). This results in 16 model-format combinations. Each model will answer every article 3 times per file-format condition, with outcomes derived from aggregated run-level results. The primary outcome is the proportion of correctly answered questions; the secondary outcome is the pass/fail rate.

RESULTS: The study has been approved by the Witten/Herdecke University Ethics Committee (S-260/2025; dated August 10, 2025) and is preregistered at the Open Science Framework. The study is supported by internal departmental resources only, and no external funding was received. Because this protocol evaluates LLMs using expired CME materials, no human participants are being recruited. Data collection is planned to begin in June 2026 and is expected to last approximately 4 weeks. At the time of manuscript submission, no data have been collected or analyzed. Results are expected to be available after the completion of data collection and statistical analysis in 2026. The analyses will quantify performance differences across document formats; these findings may inform the feasibility of nonsearchable document formats as a temporary measure to reduce LLM-enabled cheating risks in CME contexts.

CONCLUSIONS: By quantifying how document format constrains LLM performance, this study aims to evaluate simple technical safeguards that may reduce artificial intelligence-assisted manipulation of CME tests and inform regulators and CME providers about how to balance assessment validity, accessibility, and responsible LLM integration into postgraduate medical education.

PMID:42520225 | DOI:10.2196/91675

By Nevin Manimala

Portfolio Website for Nevin Manimala