Lancet Digit Health. 2026 Aug 25:101039. doi: 10.1016/j.landig.2026.101039. Online ahead of print.
ABSTRACT
BACKGROUND: Electroencephalogram (EEG) interpretation is essential for neurological diagnosis, but expert interpretation is limited globally, and existing AI methods address narrow tasks. This study aimed to develop and externally validate a broadly applicable foundation model capable of expert-level performance across diverse EEG tasks and clinical settings.
METHODS: In this multicentre study we developed a multidomain omnibus for reading and generalising over thorough EEG interpretation (MORGOTH), a foundation model that supports broad clinical interpretation across all major settings. We developed MORGOTH using EEGs from 18 677 patients across Massachusetts General Hospital, Brigham and Women’s Hospital, Beth Israel Deaconess Medical Center, and Boston Children’s Hospital, collected between Jan 1, 2003, and Feb 1, 2025, and validated it internally on 13 334 patients and externally on 1573 patients from 48 institutions, spanning diverse clinical settings and ages (0 years to >90 years). Test datasets annotated by six to 30 experts enabled inter-rater reliability (IRR) analysis by comparing model-expert and expert-expert agreement. MORGOTH was compared against both human experts and state-of-the-art models using area under the curve (AUC) and the percentage of experts’ operating points under the curve (EUC) for receiver operating characteristic (ROC) and precision-recall curves, as well as IRR and statistical calibration.
FINDINGS: MORGOTH achieved expert-level performance with AUC-ROC scores of 0·86 to 0·98 across 17 EEG findings. MORGOTH outperformed at least 90% of experts on three of seven multi-expert-annotated datasets and exceeded at least 20% of experts on each of the 17 tasks. Event-level performance was especially strong for seizure and ictal-interictal-injury continuum detection (EUC=96·6%) and spike detection (EUC=100%). IRR analysis showed that MORGOTH matched or exceeded expert consensus. External validation confirmed consistent performance with modest declines from internal to external test sets (event-level AUC -1·21%, EUC -3·33%; EEG-level AUC -2·12%, and EUC -9·52%). MORGOTH performance also remained robust across age, sex, and moderate channel loss, with lower age sensitivity (20·90% vs 30·90%) and fewer sex-related differences (33·33% vs 44·00%) than SPaRCNet.
INTERPRETATION: MORGOTH advances automated EEG interpretation with expert-level performance across clinical settings, offering improved diagnostic accuracy in low-resource environments and greater efficiency in high-volume centres.
FUNDING: US National Institutes of Health.
PMID:42642264 | DOI:10.1016/j.landig.2026.101039