scHinter
scHinter imputes dropout events in single-cell RNA sequencing (scRNA-seq) data to recover missing gene expression values with an emphasis on limited or imbalanced sample sizes.
Key Features:
- Dropout imputation: Recovers lost gene expression measurements in scRNA-seq datasets to produce an imputed expression matrix.
- Small/imbalanced sample handling: Targets datasets with limited or imbalanced numbers of cells to stabilize imputation under constrained sample sizes.
- Voting-based consensus distance: Employs a voting-based ensemble distance mechanism to enhance imputation accuracy by aggregating multiple distance perspectives.
- SMOTE (Synthetic Minority Over-sampling Technique): Uses SMOTE for random interpolation to generate synthetic data points that support imputation in sparse sample regimes.
- Hierarchical framework: Integrates a hierarchical structure to improve the reliability of imputations across different levels of analysis.
- Matlab implementation: Provided as a Matlab package for execution within Matlab-based analysis environments.
- Benchmarking results: Demonstrated superior and consistent performance across diverse scRNA-seq datasets with imbalanced or limited sample sizes relative to MAGIC, scImpute, SAVER, and netSmooth.
- Downstream compatibility: Produces imputed expression matrices intended for integration into downstream analyses such as cell type clustering, dimension reduction, and visualization.
Scientific Applications:
- Recovery of gene expression: Restores missing transcript measurements to improve the fidelity of single-cell expression profiles.
- Analysis of small or imbalanced scRNA-seq cohorts: Enables more robust analysis when cell counts are limited or class distributions are uneven.
- Downstream single-cell workflows: Supplies imputed matrices for cell type clustering, dimensionality reduction, and visualization tasks.
- Method comparison and benchmarking: Serves as a method for comparison against existing imputation approaches such as MAGIC, scImpute, SAVER, and netSmooth.
Methodology:
scHinter uses a three-module computational approach comprising a voting-based ensemble consensus distance, Synthetic Minority Over-sampling Technique (SMOTE) for random interpolation, and a hierarchical framework to improve imputation reliability in small or imbalanced scRNA-seq datasets.
Topics
Details
- Programming Languages:
- MATLAB
- Added:
- 11/14/2019
- Last Updated:
- 12/17/2020
Operations
Publications
Ye P, Ye W, Ye C, Li S, Ye L, Ji G, Wu X. scHinter: imputing dropout events for single-cell RNA-seq data with limited sample size. Bioinformatics. 2019;36(3):789-797. doi:10.1093/bioinformatics/btz627. PMID:31392316.
PMID: 31392316
Funding: - National Natural Science Foundation of China: 61573296, 61802323, 61871463
- Natural Science Foundation of Fujian Province of China: 2017J01068