Single-cell RNA-sequencing (scRNA-seq) now routinely profiles thousands of individual transcriptomes, yet no consensus pipeline exists for either laboratory protocol or downstream analysis. Droplet-based datasets are especially challenging: their count matrices are extremely sparse and many biologically important genes are expressed at very low levels. Standard workflows, borrowed from bulk RNA-seq, apply normalization, log-transformation, and feature filtering, steps that can distort true-zero counts and systematically discard low-expression regulators. A method that natively handles sparsity, avoids imputation, and retains informative low-abundance genes is therefore desirable. We present an extended version of the COTAN workflow, a framework based on gene correlations that models zero counts directly, requires no log-normalization or scaling, and is tailored to UMI droplet data. COTAN’s key features are (i) robust gene–gene correlation matrices, validated on diverse public datasets, (ii) a superior detection of differentially expressed genes, including those with very low expression, consistently outperforming leading tools, (iii) a highly sensitive statistical scoring, able to detect heterogeneous cell groups and to confirm whether a cluster is truly transcriptomically uniform. Together, these advances provide a biologically grounded, end-to-end solution for scRNA-seq analysis. COTAN is available in Bioconductor and on GitHub. Supplementary information is available at this site.
From gene correlations to cell clusters: COTAN improved scRNA-seq analysis
Galfre S. G.;Fantozzi M.;Sirbu A.;Testa I.;Tolloso M.;Priami C.;Morandin F.
2026-01-01
Abstract
Single-cell RNA-sequencing (scRNA-seq) now routinely profiles thousands of individual transcriptomes, yet no consensus pipeline exists for either laboratory protocol or downstream analysis. Droplet-based datasets are especially challenging: their count matrices are extremely sparse and many biologically important genes are expressed at very low levels. Standard workflows, borrowed from bulk RNA-seq, apply normalization, log-transformation, and feature filtering, steps that can distort true-zero counts and systematically discard low-expression regulators. A method that natively handles sparsity, avoids imputation, and retains informative low-abundance genes is therefore desirable. We present an extended version of the COTAN workflow, a framework based on gene correlations that models zero counts directly, requires no log-normalization or scaling, and is tailored to UMI droplet data. COTAN’s key features are (i) robust gene–gene correlation matrices, validated on diverse public datasets, (ii) a superior detection of differentially expressed genes, including those with very low expression, consistently outperforming leading tools, (iii) a highly sensitive statistical scoring, able to detect heterogeneous cell groups and to confirm whether a cluster is truly transcriptomically uniform. Together, these advances provide a biologically grounded, end-to-end solution for scRNA-seq analysis. COTAN is available in Bioconductor and on GitHub. Supplementary information is available at this site.I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.


