Categories
Nevin Manimala Statistics

Radiomics and machine learning analysis of ultrasound images in Bethesda IV thyroid nodules: a retrospective monocentric study

Updates Surg. 2026 Jul 31. doi: 10.1007/s13304-026-02781-w. Online ahead of print.

ABSTRACT

Bethesda IV thyroid nodules remain a major diagnostic challenge because cytology cannot reliably distinguish benign from malignant follicular-patterned lesions, often leading to diagnostic surgery for ultimately benign disease. This study evaluated whether ultrasound radiomics combined with machine learning could improve preoperative risk stratification in this setting. We conducted a retrospective monocentric study including 69 surgically treated patients with Bethesda IV thyroid nodules and definitive histopathology. Ultrasound images acquired between 2019 and 2020 were manually segmented, and radiomic features were extracted using LIFEx software. After preprocessing and removal of non-informative features, the dataset was split into training (n = 48) and a held-out test set (n = 21), isolated prior to any modelling procedure. Dimensionality reduction was performed with principal component analysis analysis (20 components, 99.5% explained variance). Three supervised machine-learning models were developed and compared: K-nearest neighbors (KNN), Random Forest (RF), and Extreme Gradient Boosting (XGBoost). Internal validation was performed by leave-one-out cross-validation on the training set, with StandardScaler and PCA fitted inside each fold. Final performance was assessed once on the held-out test set. All metrics are reported with 95% confidence intervals. Model performance was assessed using cross-validation and leave-one-out cross-validation, with evaluation of accuracy, F1-score, sensitivity, specificity, and ROC-AUC. Final histology showed malignancy in 39/69 nodules (56.5%). On the held-out test set, RF (n = 50 estimators, depth = 10) achieved the best overall performance: AUC 0.735 (95% CI 0.513-0.956), sensitivity 0.714 (95% CI 0.454-0.883), specificity 0.571 (95% CI 0.250-0.842), and accuracy 0.667 (95% CI 0.454-0.828). KNN (k = 3) achieved AUC 0.602 (95% CI 0.359-0.845). XGBoost did not generalise to the test set (AUC 0.378, 95% CI 0.094-0.661). Leave-one-out cross-validation estimates were modest across all models (AUC range 0.397-0.588), with confidence intervals overlapping 0.50, reflecting the limited statistical power of the training set at n = 48. Ultrasound radiomics combined with Random Forest shows preliminary discriminative ability for distinguishing benign from malignant Bethesda IV thyroid nodules. The wide confidence intervals observed highlight the need for a larger prospective validation cohort. These findings establish a methodological framework and provide sample size benchmarks for a powered confirmatory study.

PMID:42536308 | DOI:10.1007/s13304-026-02781-w

By Nevin Manimala

Portfolio Website for Nevin Manimala