Adv Med Educ Pract. 2026 Aug 4;17:621664. doi: 10.2147/AMEP.S621664. eCollection 2026.
ABSTRACT
BACKGROUND: Large language models (LLMs) are increasingly used as learning resources in medical education, yet their performance and reliability in microbiology, a discipline with a broad, heterogeneous knowledge base, have not been systematically evaluated across multiple platforms.
OBJECTIVE: To benchmark seven publicly available LLMs on microbiology multiple-choice questions (MCQs), assessing overall accuracy, test-retest reliability, topic-specific performance, and the relationship between cognitive complexity and model performance.
METHODS: Seven LLMs (Claude 4.6 Sonnet, Gemini 3.0, ChatGPT-5.2, Grok 4, Copilot, DeepSeek V3, and Kimi K2) completed 200 MCQs distributed across 20 microbiology topics and five Bloom’s taxonomy levels in three independent sessions separated by 24-hour intervals. A total of 4200 responses were analyzed. Statistical analysis included one-way ANOVA with Tukey’s HSD post hoc tests, repeated-measures ANOVA, intraclass correlation coefficients (ICCs), and Pearson correlations.
RESULTS: The collective mean accuracy was 86.18%. Six of seven systems exceeded the 80% high-competency threshold; Claude (89.83%), Grok (89.50%), and GPT (88.83%) led the group. Gemini (72.33%) was the only underperforming system. Test-retest reliability varied dramatically: Claude achieved excellent ICC (0.966), while Gemini exhibited poor reliability (ICC = 0.290), with session-to-session fluctuations of up to 100 percentage points on individual topics. Microbial Cell (100%) was the easiest topic; Viral Genomics (61.9%) was the most challenging across all systems. A uniform decline at Bloom’s Level 4 (Analyze) was observed across all LLMs, with no model exceeding 78%.
CONCLUSION: Contemporary LLMs demonstrate substantial knowledge of microbiology but differ markedly in reliability. Response consistency, alongside accuracy, should be a primary criterion for educational deployment. These findings are specific to microbiology MCQ performance and may not generalize to open-ended clinical reasoning.
PMID:42572759 | PMC:PMC13453379 | DOI:10.2147/AMEP.S621664