Categories
Nevin Manimala Statistics

scRecover: Discriminating True and False Zeros in Single-Cell RNA-Seq Data for Imputation

Stat Med. 2025 Feb 28;44(5):e10334. doi: 10.1002/sim.10334.

ABSTRACT

High-throughput single-cell RNA-seq (scRNA-seq) data contains an excess of zero values, which can be contributed by unexpressed genes and detection signal dropouts. Existing imputation methods fail to distinguish between these two types of zeros. In this study, we introduce a statistical framework that effectively differentiates true zeros (lack of expression) from false zeros (dropouts). By focusing only on imputing the dropout zeros, we developed a new imputation tool, scRecover. Our approach utilizes a zero-inflated negative binomial framework to model the gene expression of each gene in each cell, enabling the estimation of zero-dropout probability. Additionally, we employ a modified version of the Good and Toulmin model to identify true zeros for each gene. To achieve imputation, scRecover is combined with other imputation methods such as scImpute, SAVER and MAGIC. Down-sampling experiments show that it recovers dropout zeros with higher accuracy and avoids over-imputing true zero values. Experiments conducted on real world data highlight the ability of scRecover to enhance downstream analysis and visualization.

PMID:39912305 | DOI:10.1002/sim.10334

By Nevin Manimala

Portfolio Website for Nevin Manimala