Dino

Dino models and normalizes gene-level single-cell RNA-seq counts by fitting a flexible negative-binomial mixture to adjust full gene expression distributions for library size effects and high zero abundance, including UMI data.


Key Features:

  • Negative-binomial mixture modeling: Fits a flexible negative-binomial mixture model to gene expression counts.
  • Distribution-level normalization: Normalizes entire gene-specific expression distributions rather than only adjusting mean expression for library size.
  • Library size variance correction: Addresses non-constant variance and distributional distortions across different library sizes.
  • Zero-count accommodation: Handles high proportions of zero counts typical in scRNA-seq datasets.
  • Robustness to shallow sequencing and heterogeneity: Maintains performance under shallow sequencing depths and sample heterogeneity.
  • Mitigation of technical artifacts: Reduces technical and experimental artifacts observed even when unique molecular identifiers (UMIs) are employed.
  • Improved downstream performance: Increases statistical power and reduces false discoveries in downstream analyses.
  • Implementation: Provided as an R package.
  • Validation: Demonstrated effectiveness using simulated experiments and case-study datasets.

Scientific Applications:

  • scRNA-seq preprocessing: Normalization of single-cell RNA-seq data to provide distributionally consistent inputs for downstream analyses.
  • Datasets with shallow sequencing or high zeros: Normalization for datasets with shallow sequencing depth, heterogeneous samples, or varying proportions of zero counts.
  • Improving downstream inference: Enhancing power and reliability of downstream analyses by reducing distributional artifacts and false discoveries.

Methodology:

Fits a flexible negative-binomial mixture model to gene-level scRNA-seq counts to normalize entire gene expression distributions across library sizes and was validated using simulated experiments and case-study datasets.

Topics

Details

Cost:
Free of charge
Tool Type:
library
Programming Languages:
R
Added:
11/3/2021
Last Updated:
11/24/2024

Operations

Publications

Brown J, Ni Z, Mohanty C, Bacher R, Kendziorski C. Normalization by distributional resampling of high throughput single-cell RNA-sequencing data. Bioinformatics. 2021;37(22):4123-4128. doi:10.1093/bioinformatics/btab450. PMID:34146085. PMCID:PMC9502161.

PMID: 34146085
PMCID: PMC9502161
Funding: - National Library of Medicine Bio-Data Science Training program: T32LM012413 - National Institutes of Health: NIHGM102756

Links