Single-cell RNA-sequencing (scRNA-seq) now routinely profiles thousands of individual transcriptomes, yet no consensus pipeline exists for either laboratory protocol or downstream analysis. Droplet-based datasets are especially challenging: their count matrices are extremely sparse and many biologically important genes are expressed at very low levels. Standard workflows, borrowed from bulk RNA-seq, apply normalization, log-transformation, and feature filtering, steps that can distort true-zero counts and systematically discard low-expression regulators. A method that natively handles sparsity, avoids imputation, and retains informative low-abundance genes is therefore desirable. We present an extended version of the COTAN workflow, a framework based on gene correlations that models zero counts directly, requires no log-normalization or scaling, and is tailored to UMI droplet data. COTAN’s key features are (i) robust gene–gene correlation matrices, validated on diverse public datasets, (ii) a superior detection of differentially expressed genes, including those with very low expression, consistently outperforming leading tools, (iii) a highly sensitive statistical scoring, able to detect heterogeneous cell groups and to confirm whether a cluster is truly transcriptomically uniform. Together, these advances provide a biologically grounded, end-to-end solution for scRNA-seq analysis. COTAN is available in Bioconductor and on GitHub. Supplementary information is available at this site.

From gene correlations to cell clusters: COTAN improved scRNA-seq analysis

Galfre S. G.;Fantozzi M.;Sirbu A.;Testa I.;Tolloso M.;Priami C.;Morandin F.
2026-01-01

Abstract

Single-cell RNA-sequencing (scRNA-seq) now routinely profiles thousands of individual transcriptomes, yet no consensus pipeline exists for either laboratory protocol or downstream analysis. Droplet-based datasets are especially challenging: their count matrices are extremely sparse and many biologically important genes are expressed at very low levels. Standard workflows, borrowed from bulk RNA-seq, apply normalization, log-transformation, and feature filtering, steps that can distort true-zero counts and systematically discard low-expression regulators. A method that natively handles sparsity, avoids imputation, and retains informative low-abundance genes is therefore desirable. We present an extended version of the COTAN workflow, a framework based on gene correlations that models zero counts directly, requires no log-normalization or scaling, and is tailored to UMI droplet data. COTAN’s key features are (i) robust gene–gene correlation matrices, validated on diverse public datasets, (ii) a superior detection of differentially expressed genes, including those with very low expression, consistently outperforming leading tools, (iii) a highly sensitive statistical scoring, able to detect heterogeneous cell groups and to confirm whether a cluster is truly transcriptomically uniform. Together, these advances provide a biologically grounded, end-to-end solution for scRNA-seq analysis. COTAN is available in Bioconductor and on GitHub. Supplementary information is available at this site.
2026
Galfre, S. G.; Fantozzi, M.; Sirbu, A.; Testa, I.; Tolloso, M.; Alberti, A.; Priami, C.; Morandin, F.
File in questo prodotto:
Non ci sono file associati a questo prodotto.

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11568/1366388
 Attenzione

Attenzione! I dati visualizzati non sono stati sottoposti a validazione da parte dell'ateneo

Citazioni
  • ???jsp.display-item.citation.pmc??? 1
  • Scopus 0
  • ???jsp.display-item.citation.isi??? 0
social impact