Dino
Dino models and normalizes gene-level single-cell RNA-seq counts by fitting a flexible negative-binomial mixture to adjust full gene expression distributions for library size effects and high zero abundance, including UMI data.
Key Features:
- Negative-binomial mixture modeling: Fits a flexible negative-binomial mixture model to gene expression counts.
- Distribution-level normalization: Normalizes entire gene-specific expression distributions rather than only adjusting mean expression for library size.
- Library size variance correction: Addresses non-constant variance and distributional distortions across different library sizes.
- Zero-count accommodation: Handles high proportions of zero counts typical in scRNA-seq datasets.
- Robustness to shallow sequencing and heterogeneity: Maintains performance under shallow sequencing depths and sample heterogeneity.
- Mitigation of technical artifacts: Reduces technical and experimental artifacts observed even when unique molecular identifiers (UMIs) are employed.
- Improved downstream performance: Increases statistical power and reduces false discoveries in downstream analyses.
- Implementation: Provided as an R package.
- Validation: Demonstrated effectiveness using simulated experiments and case-study datasets.
Scientific Applications:
- scRNA-seq preprocessing: Normalization of single-cell RNA-seq data to provide distributionally consistent inputs for downstream analyses.
- Datasets with shallow sequencing or high zeros: Normalization for datasets with shallow sequencing depth, heterogeneous samples, or varying proportions of zero counts.
- Improving downstream inference: Enhancing power and reliability of downstream analyses by reducing distributional artifacts and false discoveries.
Methodology:
Fits a flexible negative-binomial mixture model to gene-level scRNA-seq counts to normalize entire gene expression distributions across library sizes and was validated using simulated experiments and case-study datasets.
Topics
Details
- Cost:
- Free of charge
- Tool Type:
- library
- Programming Languages:
- R
- Added:
- 11/3/2021
- Last Updated:
- 11/24/2024
Operations
Publications
Brown J, Ni Z, Mohanty C, Bacher R, Kendziorski C. Normalization by distributional resampling of high throughput single-cell RNA-sequencing data. Bioinformatics. 2021;37(22):4123-4128. doi:10.1093/bioinformatics/btab450. PMID:34146085. PMCID:PMC9502161.
PMID: 34146085
PMCID: PMC9502161
Funding: - National Library of Medicine Bio-Data Science Training program: T32LM012413
- National Institutes of Health: NIHGM102756
Links
Repository
https://zenodo.org/record/4897558