Med Image Anal. 2026 Jun 16;113:104159. doi: 10.1016/j.media.2026.104159. Online ahead of print.
ABSTRACT
Despite high performance of deep learning in medical imaging applications, the critical lack of validated computational metrics for explainable AI (XAI) impedes clinical integration. To address this gap, our study introduces a multi-level validation framework to rigorously assess seven computational evaluation metrics applied to eight widely-used post-hoc attribution methods – spanning gradient-based, input attribution, and decomposition-based families – on large-scale structural MRI datasets (UK Biobank and ADNI) with two different benchmark tasks across several deep learning architectures. At the first level, we benchmark metrics against three scales of clinical relevance: voxel-based morphometry, regional volumetric associations, and expert radiologist assessments, revealing that many widely used conventional metrics exhibit negligible correlation with clinical evidence. At the second level, we show that metrics in addition are often significantly confounded by model architecture rather than reflecting explanation quality. At the third level, we conduct a computational efficiency and stability analysis focused on practical dimensions of the metrics. Across all levels, our statistical analyses show that only one of the metrics has consistent high performance: the Robustness score – defined as the stability of explanations across random training initializations – demonstrates strong alignment with morphometry, volumetric associations, and human expert judgments, effectively isolates the quality of the XAI method from architectural bias, and possesses good computational efficiency. Our novel, multi-level framework therefore establishes Robustness as a superior, scalable proxy for clinical validity and trustworthy AI in medical imaging.
PMID:42475759 | DOI:10.1016/j.media.2026.104159