SAME-clustering

SAME-clustering performs ensemble clustering of single-cell RNA-seq (scRNA-seq) data by aggregating multiple clustering solutions using a mixture model to improve cluster assignments and estimate the number of clusters.


Key Features:

  • Mixture model-based ensemble: Aggregates multiple clustering outputs via a mixture model to produce a consensus clustering solution.
  • Diverse-subset selection: Evaluates individual clustering solutions and selects a maximally diverse subset for combination.
  • Improved clustering accuracy: Demonstrated superior clustering performance compared to standalone methods across evaluated datasets.
  • Benchmarking across datasets: Validated on 15 scRNA-seq datasets with cluster counts ranging from 3 to 15 and cell counts from 49 to 32,695.
  • Platform robustness: Applied successfully to datasets generated by various scRNA-seq platforms.
  • Broad applicability: Ensemble methodology is applicable to clustering problems beyond scRNA-seq.

Scientific Applications:

  • Cell type identification: Identifies distinct cell types and their transcriptomic signatures within scRNA-seq data to elucidate tissue complexity.
  • Characterization of cellular heterogeneity: Supports analysis of cellular heterogeneity and dynamic processes, including investigations relevant to disease mechanisms.
  • General clustering tasks: Provides an ensemble framework applicable to diverse clustering challenges beyond single-cell transcriptomics.

Methodology:

SAME-clustering generates individual clustering solutions using different methods, evaluates solution diversity to select a maximally diverse subset, and combines the selected solutions via a mixture model ensemble.

Topics

Details

Programming Languages:
R, C++
Added:
1/14/2020
Last Updated:
11/24/2024

Operations

Publications

Huh R, Yang Y, Jiang Y, Shen Y, Li Y. SAME-clustering: Single-cell Aggregated Clustering via Mixture Model Ensemble. Nucleic Acids Research. 2019;48(1):86-95. doi:10.1093/nar/gkz959. PMID:31777938. PMCID:PMC6943136.

PMID: 31777938
PMCID: PMC6943136
Funding: - National Institutes of Health: R01HG006292, R01HL129132