Categories
Nevin Manimala Statistics

Instability of LLM text embeddings for unsupervised dimension reduction of tabular data

Front Bioinform. 2026 Jul 27;6:1851023. doi: 10.3389/fbinf.2026.1851023. eCollection 2026.

ABSTRACT

Large language model (LLM) text embeddings have recently been used for supervised learning on tabular data by serializing each observation into text and then converting the text into a dense fixed-length vector. This strategy is appealing for biomedical tabular data, which are often mixed-type and contain missing values, because it produces a complete numeric representation even when some original entries are missing. However, its suitability for unsupervised tasks remains unclear. Here we evaluate the use of LLM-derived text embeddings for dimension reduction of tabular data, focusing on biological and clinical datasets. We compare an LLM embedding-based approach with a direct tabular approach that computes dissimilarities directly from the original variables. Because unsupervised dimension reduction has no ground-truth low-dimensional target, we assess performance through stability under perturbation. Across multiple datasets and analysis settings, the LLM embedding-based approach is consistently less stable than the direct tabular approach. In particular, small amounts of additional missingness and random permutation of feature order can substantially alter the resulting low-dimensional representation. These results suggest that the straightforward use of LLM text embeddings is not reliable for unsupervised dimension reduction of tabular data.

PMID:42577418 | PMC:PMC13454057 | DOI:10.3389/fbinf.2026.1851023

By Nevin Manimala

Portfolio Website for Nevin Manimala