scAIDE

scAIDE implements an unsupervised deep learning clustering framework to identify and categorize putative and rare cell types and infer developmental trajectories from large-scale single-cell RNA-seq datasets.


Key Features:

  • Unsupervised deep learning framework: Integrates an autoencoder-imputation network with a distance-preserved embedding network (AIDE) to learn robust representations from noisy single-cell RNA-seq data.
  • Scalability and efficiency: Processes millions of cells and demonstrated analysis of 1.3 million neural cells in 30 minutes.
  • Random projection hashing-based k-means: Employs a random projection hashing-based k-means clustering algorithm to enhance detection of rare cell types.
  • Identification of cell types and trajectories: Detected 64 clusters mapped to 19 putative cell types and revealed three developmental trajectories among neural stem cells in the neural dataset.
  • Application in disease research: Classified subpopulations of malignant cells in a glioblastoma dataset.

Scientific Applications:

  • Cellular heterogeneity and development: Characterizing cellular heterogeneity and developmental trajectories across complex tissues using single-cell RNA-seq.
  • Rare cell type discovery: Detecting and profiling rare cell types within large cell populations.
  • Cancer and translational research: Classifying malignant cell subpopulations in cancers such as glioblastoma to support translational studies and personalized medicine.

Methodology:

Combines an autoencoder-imputation network with a distance-preserved embedding network (AIDE) and applies random projection hashing-based k-means clustering for unsupervised identification of cell clusters from single-cell RNA-seq data.

Topics

Details

Tool Type:
command-line tool, library
Programming Languages:
Python, R
Added:
3/19/2021
Last Updated:
4/3/2021

Operations

Publications

Xie K, Huang Y, Zeng F, Liu Z, Chen T. scAIDE: clustering of large-scale single-cell RNA-seq data reveals putative and rare cell types. NAR Genomics and Bioinformatics. 2020;2(4). doi:10.1093/nargab/lqaa082. PMID:33575628. PMCID:PMC7671411.

PMID: 33575628
PMCID: PMC7671411
Funding: - National Natural Science Foundation of China: 61673241, 61721003, 61872218, 61906105 - National Key Research and Development Program of China: 2019YFB1404804