Categories
Nevin Manimala Statistics

Novel nested conformal prediction analysis to unravel complexity in patient subtyping

Front Artif Intell. 2026 Jul 20;9:1844254. doi: 10.3389/frai.2026.1844254. eCollection 2026.

ABSTRACT

Patient subtyping is significantly challenged by intra-sample heterogeneity, which limits the effectiveness of traditional multi-class classification approaches enforcing mutually exclusive labels. Despite recent promising results in the transition from multi-class to multi-label classification, this process is not straightforward and proves hard to systemize, especially when working with datasets of small dimensions. Here, we design a novel approach, implemented in a computational framework, leveraging machine learning models and an original nested conformal prediction strategy for small datasets to enable a rigorous transition to multi-label classification. Conformal prediction can indeed offer a statistically sound approach for robust subtype predictions; yet, so far, it has been limited to be exploited only in large datasets, not often available in real biomedical scenarios. We evaluate our approach using breast cancer patient RNA-seq data from The Cancer Genome Atlas, with subtypes originally defined via the PAM50 classifier. From a methodological standpoint, our innovative nested strategy enables the successful application of conformal prediction, which commonly requires large datasets, to a relatively small sample size. Specifically, conformal prediction is adapted and applied both to the reference PAM50 classification and to multiple machine learning models to test coherence and compare the obtained multi-label predictions. Our results reveal subtype- and sample-specific complexity and highlight consistent patterns of ambiguity across models. Importantly, we distinguish between robust, inherently heterogeneous, and weak assignments, providing deeper insight into clear subtype confirmations and complexity, but also critical incoherent predictions. Thus, overall, this work demonstrates the utility of our framework for uncertainty-aware sample subtyping, particularly in small and highly unbalanced sample settings.

PMID:42548800 | PMC:PMC13429844 | DOI:10.3389/frai.2026.1844254

By Nevin Manimala

Portfolio Website for Nevin Manimala