Categories
Nevin Manimala Statistics

Multinational Validation of a Radiography-Based AI Tool for Appendicular Fracture Detection and Corresponding Reader Performance

Radiology. 2026 Aug;320(2):e243548. doi: 10.1148/radiol.243548.

ABSTRACT

Background Artificial intelligence (AI) models for fracture detection have shown promise in improving diagnostic efficiency and accuracy in radiology workflows. However, despite potential differences in imaging protocols and clinical workflows in different health care systems and settings, the performance across countries of these models and their effects on reader performance remain underexplored. Purpose To evaluate the performance of a pretrained fracture detection tool across three European hospitals and assess its impact on clinician performance. Materials and Methods In this retrospective, multireader, multicase crossover study, the radiographs of patients aged 21 years or older with suspected appendicular fractures acquired from April 2018 to July 2022, across three European hospitals, were analyzed. Reference standards were established by expert local radiologists. AI tool performance was evaluated with the area under the receiver operator characteristic curve, sensitivity, and specificity. Sensitivity and specificity were statistically compared across hospitals and per-anatomy subgroup with the χ2 test and across readers with generalized estimating equations. Results A total of 1500 patients (mean age, 43.3 years ± 19.5 [SD]; 820 men) were included. The hospital-level sensitivities of the AI tool ranged from 400 of 500 (80%; 95% CI: 76, 84) to 430 of 500 (86%; 95% CI: 83, 90) and specificities ranged from 465 of 500 (93%; 95% CI: 91, 95) to 480 of 500 (96%; 95% CI: 95, 98). Anatomy-based subgroup analysis revealed lower sensitivity of 105 of 140 (75%) for wrist/hand and finger examinations in one hospital versus 125 of 138 (91%) and 150 of 160 (94%) (P = .004), but no differences in overall sensitivity or specificity were observed across all hospitals (P = .23 and P = .10, respectively). Use of the AI tool improved overall reader sensitivity (11 percentage points [95% CI: 6.8, 14.4; P < .001]) while maintaining specificity (0.6 percentage points [95% CI: -0.6, 1.7; P = .34]). Conclusion No differences in AI performance were observed across centers, and AI assistance was associated with higher reader sensitivity. © The Author(s) 2026. Published by the Radiological Society of North America under a CC BY 4.0 license. Supplemental material is available for this article See also the editorial by Maas and Beekman in this issue.

PMID:42610815 | DOI:10.1148/radiol.243548

By Nevin Manimala

Portfolio Website for Nevin Manimala