JMIR Form Res. 2026 Aug 21;10:e94172. doi: 10.2196/94172.
ABSTRACT
BACKGROUND: Cesarean section (C-section) is the most common surgical procedure in the United States, yet its use varies widely across regions and institutions. Although clinical risk factors are central to delivery decisions, geographic context, health system capacity, and local practice patterns may also influence C-section use. Understanding both the determinants and predictability of C-section delivery is important for improving obstetric quality and equity.
OBJECTIVE: This study aims to document geographic variation in C-section use across the United States, identify maternal and county-level factors associated with C-section delivery, and evaluate the predictive performance of machine learning models across clinically defined risk groups.
METHODS: This population-based study used 38,133,279 US births from the 2013-2022 National Vital Statistics System Natality Detailed Files. County identifiers were linked to national county-level measures of insurance coverage, health care capacity, and socioeconomic conditions. Logistic regression models with county fixed effects were used for feature interpretation, and supervised machine learning models were used for prediction. Analyses were conducted separately for the full sample, a low-risk sample (n=17,760,772), and a high-risk sample (n=20,372,438). Predictive performance was evaluated using accuracy, precision, recall, F1-score, and area under the receiver operating characteristic curve (AUC), with 200-bootstrap 95% CIs. Temporal validation was also conducted using training on earlier years and testing on later years.
RESULTS: County-level C-section rates declined modestly from 32.35% in 2013 to 31.49% in 2022, but substantial geographic variation persisted, with consistently higher rates in the US South. In the full sample, model discrimination was good, with AUC values ranging from 0.8310 for logistic regression to 0.8401 for extreme gradient boosting (XGBoost). Predictive performance was substantially weaker in the low-risk sample (AUC 0.7246-0.7410) than in the high-risk sample (AUC 0.8404-0.8568). In the high-risk sample, XGBoost achieved the highest AUC (0.8568) and F1-score (0.7579), while random forest achieved the highest recall (0.7142). Temporal validation yielded similar results in the full sample (AUC 0.8335-0.8387), temporal low-risk sample (AUC 0.7364-0.7432), and temporal high-risk sample (AUC 0.8360-0.8479), indicating stable performance over time.
CONCLUSIONS: C-section use in the United States is shaped by both maternal clinical risk and geographic context. Machine learning models perform well overall and especially well in high-risk pregnancies, but prediction is substantially more difficult in low-risk pregnancies, where discretionary and contextual influences may play a larger role. These findings support the use of risk-adjusted, context-aware prediction tools for audit, benchmarking, and clinical decision support while underscoring the need for cautious implementation, subgroup monitoring, and further external validation.
PMID:42628029 | DOI:10.2196/94172