Categories
Nevin Manimala Statistics

Multiclass diagnostic performance and error patterns of a multimodal large language model in oral histopathology: A WHO-aligned study

J Oral Biosci. 2026 Jul 30;68(5):100802. doi: 10.1016/j.job.2026.100802. Online ahead of print.

ABSTRACT

OBJECTIVES: To evaluate the diagnostic performance, feature-level agreement, and error patterns of a multimodal large language model (MLLM) across the oral dysplasia-malignancy continuum.

METHODS: In this retrospective diagnostic accuracy study, 300 histopathology images representing normal oral mucosa (n = 100), oral potentially malignant disorders (OPMD) (n = 100), and oral squamous cell carcinoma (OSCC) (n = 100) were evaluated using a standardized two-stage prompting framework consisting of morphological feature extraction followed by structured diagnostic classification. Reference diagnoses were established by three experienced oral pathologists. Interobserver agreement among pathologists was assessed using Fleiss’ κ statistics. Primary outcomes included overall accuracy, macro-F1 score, and balanced accuracy. Secondary outcomes included feature-level agreement, confidence-accuracy alignment, and clinically significant error patterns.

RESULTS: Interobserver agreement among the reference pathologists was substantial for the primary three-class diagnosis (Fleiss’ κ = 0.81). Overall accuracy was 0.667 (95% confidence interval [CI]: 0.620-0.727), with a macro-F1 score of 0.662 and balanced accuracy of 0.755. Performance varied substantially across diagnostic categories, with high sensitivity for normal mucosa (0.990; 95% CI: 0.968-1.000). Reduced sensitivities were observed for OPMD (0.430; 95% CI: 0.340-0.531) and OSCC (0.580; 95% CI: 0.495-0.688), indicating reduced discriminative ability relative to normal mucosa. Feature-level agreement was generally low, with the weakest agreement for architectural dysplasia features (κ = -0.116 to -0.068), slight agreement for invasion-related features (κ = 0.111-0.120), and the highest agreement was for nuclear hyperchromasia (κ = 0.281). Most misclassifications (84.8%) occurred at the dysplasia-malignancy interface. Proxy calibration analysis demonstrated systematic overconfidence in the histopathological predictions.

CONCLUSIONS: The MLLM demonstrated moderate performance in oral histopathology, with reliable recognition of normal tissue and reduced discriminative ability for OPMD and OSCC relative to normal mucosa, particularly at the dysplasia-malignancy interface. These findings support a potential assistive role under expert supervision while highlighting important limitations in clinically critical diagnostic settings.

PMID:42531661 | DOI:10.1016/j.job.2026.100802

By Nevin Manimala

Portfolio Website for Nevin Manimala