J Med Internet Res. 2026 Aug 6;28:e89963. doi: 10.2196/89963.
ABSTRACT
BACKGROUND: Large language models (LLMs) are increasingly applied in clinical decision support, yet their diagnostic performance in Chinese-language settings and under realistic clinical workflows remains unclear. In particular, how LLMs perform across diseases with different prevalence and under stepwise diagnostic processes has not been well characterized.
OBJECTIVE: This study aimed to evaluate the diagnostic capabilities of LLMs for common diseases and rare diseases using clinical vignettes within a hypothetico-deductive framework and to identify their potential and limitations for clinical diagnosis.
METHODS: We evaluated 4 Chinese LLMs (Doubao 1.5, DeepSeek-V3, Kimi K1.5, and Leftdoctor GPT 3.5) using 56 clinical cases (28 chronic obstructive pulmonary disease [COPD], and 28 relapsing polychondritis [RP]) sourced from the China Clinical Case Results Database (March 31-April 14, 2025). Patient information was provided incrementally, starting with the initial medical history, followed by physical examination, and laboratory results. Evaluation metrics included top-3 accuracy (RTop3D), top-1 accuracy (RTopD), final diagnostic accuracy (RFA), and mean reciprocal rank (MRR). Statistical analysis was performed using generalized estimating equations (GEE), Friedman tests, and Wilcoxon signed-rank tests with Bonferroni correction. In addition, a qualitative analysis was conducted to characterize recurrent patterns of diagnostic errors.
RESULTS: LLMs demonstrated significantly higher diagnostic accuracy for COPD compared to RP across all metrics (P<.001). Diagnostic accuracy improved after additional clinical information was provided, with the improvement mainly observed in RP cases. In RP, diagnostic accuracy increased from 32.14% to 71.43% for DeepSeek and from 35.71% to 78.57% for Doubao, whereas COPD accuracy remained consistently high across all diagnostic stages (82.14%-92.86%). For COPD, ranking performance was high and comparable among all models (MRR range: 0.82-0.89; P=.71). In RP, diagnostic performance differed significantly among models (MRR range: 0.10-0.39; P<.001). Qualitative analysis showed that COPD errors were mainly related to a failure to recognize specific features, whereas RP errors involved more diverse patterns, particularly the neglect of negative evidence and the failure to recognize specific features.
CONCLUSIONS: Chinese LLMs demonstrated relatively strong diagnostic performance for common diseases such as COPD, but lower and less stable performance for rare diseases such as RP. Additional clinical information improved diagnostic accuracy primarily in RP cases, although differences between models remained evident under diagnostically complex conditions. Error patterns in RP cases suggest that current LLMs remain limited in their ability to integrate complex clinical information and exclusionary findings. Careful evaluation and appropriate clinical oversight remain important for their application in clinical practice.
PMID:42561411 | DOI:10.2196/89963