Front Genet. 2026 Jul 6;17:1818099. doi: 10.3389/fgene.2026.1818099. eCollection 2026.
ABSTRACT
Understanding tissue-specific transcriptomic structures in livestock is essential for elucidating the molecular basis of economically important traits. Conventional differential gene expression analysis efficiently identifies genes with large average expression differences but does not fully capture multivariate expression structures and gene-gene interaction patterns that define tissue identity. In this study, we developed an explainable machine learning framework to classify seven Hanwoo cattle tissues using RNA sequencing data and to systematically compare the relative contributions of statistical and model-derived signals. A Random Forest-based one-versus-rest classification model was trained on 130 Hanwoo transcriptomes and externally validated using 231 independent Bos taurus samples derived from heterogeneous public datasets following reference-based batch correction. Repeated balanced validation demonstrated stable generalization performance, achieving a mean accuracy of 0.907 and a macro-average area under the receiver operating characteristic curve of 0.963. A comparative analysis of gene sets selected by differential expression analysis, genes prioritized by model-based feature attribution, their union, and randomly selected genes within the same classification framework revealed that strongly differentially expressed genes form the primary discriminatory structure for tissue classification. In contrast, integration of model-prioritized genes enhanced classification performance, particularly for biologically related tissues, whereas randomly selected genes produced reduced and unstable predictive performance. Model interpretation further revealed that highly contributory genes were consistent with known tissue-specific biological functions and exhibited non-linear, expression-dependent contribution patterns shaped by coordinated multigene contexts. These findings indicate that tissue identity is supported by a hierarchical transcriptional structure in which dominant differential signals establish primary class boundaries and multivariate interaction patterns refine decision surfaces. The proposed framework provides an interpretable strategy for distinguishing statistical significance from predictive relevance in high-dimensional transcriptomic data and offers practical implications for molecular marker development in livestock genomics and breeding programs.
PMID:42473676 | PMC:PMC13381022 | DOI:10.3389/fgene.2026.1818099