Trop Anim Health Prod. 2026 Jul 24;58(7):446. doi: 10.1007/s11250-026-05261-w.
ABSTRACT
For the first time, this study employs a century-long dataset (1925-2024) to reveal key factors that would be in relationship with sheep population (NSheep) in Türkiye using state-of-the-art machine learning algorithms. Due to the existence of missing values in the original dataset, missing observations were addressed through four imputation techniques-Next Observation Carried Backward (NOCB), Mean, MIDASpy, and Random Forest (RF)-generating four distinct datasets for comparative analysis. For revealing the key production factors related with NSheep, Extreme Gradient Boosting (XGB) and Multilayer Perceptron (MLP) algorithms were modeled via 5-fold cross-validation and multiple performance metrics (R², MSE, RMSE, MAE, and MdAPE). MLP produced lower prediction errors than XGB across all imputation techniques, though this difference was statistically confirmed only under NOCB and RF imputation (Diebold-Mariano test, P < 0.01 and P < 0.05, respectively); differences under MEAN and MIDASpy imputation were not significant. The highest overall accuracy was achieved by MLP with NOCB imputation (R² = 0.975), while XGB with RF imputation showed the weakest fit (R² = 0.917). Feature importance analyses consistently identified cattle population (NBovine) as the dominant variable associated with NSheep across all four imputation techniques and both algorithms, followed by meadow and pasture area (M&PH) for XGBoost and a more evenly distributed set of variables (M&PH, sheep meat production, goat population) for MLP. Given that NBovine and NSheep both increased steadily over the study period, this association is interpreted as reflecting shared structural growth among livestock subsectors rather than a causal effect of cattle population on sheep numbers. These dataset-specific, associative findings support the value of combining multiple imputation strategies with flexible machine learning algorithms to characterize structural interdependencies in long-term agricultural production data. Future studies incorporating chronologically ordered validation schemes and explicitly modeling structural breaks and policy shifts could further clarify the robustness of these associations, and extending this approach to other livestock species and regions would help establish their broader generalizability.
PMID:42496899 | DOI:10.1007/s11250-026-05261-w