Categories
Nevin Manimala Statistics

Are Limited Electronic Medical Record Follow-Up Data Sufficiently Useful for Validating the Performance of Survival Prediction Models?

JCO Clin Cancer Inform. 2026 Jul-Sep;10(3):e2600105. doi: 10.1200/CCI-26-00105. Epub 2026 Aug 28.

ABSTRACT

PURPOSE: Machine learning models that predict survival time for patients are increasingly used for clinical decision support. Validating model performance in deployment is important but challenging because the only timely source of follow-up/death data is the electronic medical record (EMR), which is known to undercapture deaths, resulting in informative censoring. We examined whether model performance evaluation using EMR data alone can distinguish between low- and high-quality models using high-quality cancer registry data combined with EMR data as a comparison to calculate model performance.

MATERIALS AND METHODS: This was a retrospective study of 3,330 patients with metastatic cancer diagnosed from 2008 to 2018. We used regularized Cox proportional hazards regression trained on features from the EMR to predict length of survival after diagnosis. We trained models by varying the number of features to span a range of baseline discrimination performance and validated performance first using higher-quality EMR and cancer registry outcome data as a reference standard, followed by EMR data alone to simulate the scenario when real-time validation is performed and only EMR data are available.

RESULTS: The model with all features had a C-index of 0.66 (95% CI, 0.65 to 0.68) and an integrated Brier score (IBS) of 0.17 (95% CI, 0.16 to 0.18) when validating with reference standard data, compared with a C-index of 0.67 (95% CI, 0.65 to 0.69) and an IBS of 0.16 (95% CI, 0.15 to 0.17) with EMR data only. When using fewer features, performance estimates dropped similarly in both scenarios. When using reference standard data, the model was found to be reasonably calibrated, but with EMR data, the model was incorrectly found to systematically underpredict survival.

CONCLUSION: EMR data were useful for validation of model discrimination, but not calibration.

PMID:42664478 | DOI:10.1200/CCI-26-00105

By Nevin Manimala

Portfolio Website for Nevin Manimala