Summary

A bioRxiv preprint using data from 1,494 donors finds that many single-cell gene-expression studies may lack sufficient statistical power to detect small effects. Cell counts, gene-expression levels and sequencing depth strongly influenced which findings were detected and reproduced.

A bioRxiv preprint reports that many single-cell gene-expression studies are not large or technically sensitive enough to reliably detect small differences between biological groups. The analysis used sex-biased differential expression in brain cells from 1,494 donors to estimate how often different experimental designs could identify real effects.

The authors found substantial loss of statistical power even in experiments involving up to 600 donors. They also found that the number of cells measured, the expression level of a gene and sequencing depth strongly influenced whether a gene was detected and whether the result could be reproduced.

The work is a preprint from researchers at Imperial College London, King’s College London, the Icahn School of Medicine at Mount Sinai and the University of Cambridge.

Why small effects are difficult to detect

Single-cell RNA sequencing measures gene activity separately in individual cells. It can reveal whether genes are more or less active in a particular cell type under different conditions, but the measurements contain both biological variation and technical noise.

Statistical power is the probability that an experiment will detect an effect of a given size when that effect is present. Small expression differences require more informative observations than large differences. In single-cell studies, the relevant information depends not only on the number of donors but also on how many cells of the target type are captured and how deeply their RNA is sequenced.

The researchers used sex-biased differential expression across three brain cell types as an empirical basis for power estimates. Rather than relying only on small pilot datasets or simulated measurements, they examined results from the larger donor set and evaluated how different study configurations affected discovery and reproducibility.

Sequencing depth and cell counts changed the results

In one analysis, the authors reduced astrocyte data to match the poorer sequencing characteristics of microglia. More than half of the differentially expressed genes identified in the full dataset were then lost. This result illustrates how a study can detect fewer genes not because the underlying biology has changed, but because the cell population is measured with fewer or lower-quality observations.

The analysis also found that conventional statistical significance thresholds were not enough to make all reported findings equally reproducible. After multiple-testing correction, only the top quartile of significant genes was reproducible in the study’s empirical assessment.

A predictive model fitted to the results identified cell count and a gene’s expression level as strong determinants of statistical power. The findings suggest that donor numbers alone are an incomplete way to assess whether a single-cell experiment can answer its intended question.

Implications for single-cell study design

The authors recommend enriching experiments for the cell types of interest and sequencing more deeply, particularly when studying rarer populations such as microglia. They also call for explicit power calculations before experiments begin and suggest using stricter significance thresholds when identifying new findings.

The study’s evidence is computational and statistical: it derives power estimates from an existing dataset rather than testing a treatment or measuring a clinical outcome. Its direct analysis concerns three brain cell types and sex-biased expression, so the results provide an empirical warning for comparable single-cell designs rather than a universal performance estimate for every tissue or research question.

The work is currently available as a bioRxiv preprint. Its central contribution is a practical one: a large donor count can still leave small gene-expression effects difficult to detect when the relevant cell population is scarce or sequencing is shallow.

Sources