Categories
Nevin Manimala Statistics

Identification of hidden subtypes in occupational health examinations and their biomedical characteristics using graph-enhanced deep representation learning

Front Public Health. 2026 Jul 13;14:1847173. doi: 10.3389/fpubh.2026.1847173. eCollection 2026.

ABSTRACT

PURPOSE: Occupational health examinations have long relied on conventional health classes to assess worker health status. However, individuals within the same examination class may differ substantially in risk sources, pathophysiological background, and potential progression patterns. This study aimed to determine whether stable hidden health subtypes exist beyond conventional occupational health classifications in a large real-world cohort and to evaluate the utility of graph-enhanced deep representation learning for subtype discovery and biomedical interpretation.

METHODS: We included 63,988 occupational health examination records collected from five organizational units between 2022 and 2025. A standardized master dataset was constructed by harmonizing anthropometric measures, blood pressure, glucose metabolism, lipid metabolism, liver function, kidney function, inflammatory markers, and vascular-related indicators. We first established a benchmark panel including conventional statistical models, tree-based models, and neural network models. We then developed a domain-aware prototype tabular network (DAPTN) and its graph-enhanced variant (Graph-DAPTN) to learn individual-level latent health representations, prototype distances, and hidden subtype structure. The identified subtypes were further interpreted in terms of clinical abnormality burden, risk gradients, temporal drift, and cross-domain stability.

RESULTS: The cohort exhibited substantial heterogeneity across units and years, as well as strong coupling among biomedical indicators. Linear models showed limited performance in capturing complex occupational health phenotypes, whereas nonlinear models performed better overall. In subtype-discovery validation, Graph-DAPTN showed stronger cluster separation and class-subtype concordance than several conventional baselines. Based on this representation, six clinically interpretable hidden subtypes were identified beyond conventional health classes: resilient-healthy, hepato-inflammatory, uric-metabolic, gluco-lipotoxic, hypertensive-homocysteine, and ageing-vascular. These subtypes showed distinct abnormality spectra in glucose and lipid metabolism, blood pressure, liver enzymes, uric acid, homocysteine, and inflammatory burden, and their distributions varied across years and organizational units.

CONCLUSION: Conventional health classes in occupational health examinations do not fully capture the underlying heterogeneity of worker health status. Graph-enhanced deep representation learning can improve health-state modeling while identifying clinically meaningful hidden subtypes beneath routine classification. These findings may support more refined risk stratification, key-population screening, and unit-level health intervention planning in occupational health practice.

PMID:42517104 | PMC:PMC13402512 | DOI:10.3389/fpubh.2026.1847173

By Nevin Manimala

Portfolio Website for Nevin Manimala