Categories
Nevin Manimala Statistics

Modeling Depression as a Gradient Condition Using Acoustic Features of Mandarin Speech

J Speech Lang Hear Res. 2026 Aug 6:1-15. doi: 10.1044/2026_JSLHR-26-00145. Online ahead of print.

ABSTRACT

PURPOSE: Depression is prevalent yet frequently underdiagnosed. Although speech-based detection methods show promise, most prior studies have emphasized binary classification in nontonal languages. This study examined whether acoustic features of spontaneous Mandarin speech reflect depression as a categorical condition or a continuous, dimensional construct.

METHOD: A validated emotional Mandarin speech corpus paired with self-reported depression severity scores was analyzed. Acoustic features were extracted using the extended Geneva Minimalistic Acoustic Parameter Set. A multistage analytic framework was applied, including random forest classification to distinguish depressed from nondepressed clips, linear mixed-effects modeling to examine group and severity effects, and unsupervised clustering to examine latent structure in the acoustic feature space.

RESULTS: The random forest classifier achieved an accuracy of 72.3% and an area under the receiver operating characteristic curve of 0.826 on the held-out test set. Model performance favored sensitivity over precision in identifying depressed clips. SHapley Additive exPlanations analysis identified features across multiple acoustic domains, including spectral, pitch-related, voice quality, cepstral, and intensity measures, as important contributors to model predictions. Among the Top 10 features, no acoustic features showed statistically significant group differences. In contrast, median and mean fundamental frequencies were significantly positively associated with depression severity. Clustering analysis revealed overlapping group structures and a continuous distribution of severity scores across clusters.

CONCLUSIONS: These findings provide preliminary support for a dimensional conceptualization of depression in spontaneous Mandarin speech. While classification models can distinguish depressed from nondepressed speech with moderate accuracy, acoustic features appear to vary more consistently with symptom severity than with diagnostic group. Speech-based measures may therefore have potential for continuous monitoring of depression, although this possibility requires further longitudinal verification.

PMID:42560658 | DOI:10.1044/2026_JSLHR-26-00145

By Nevin Manimala

Portfolio Website for Nevin Manimala