SUMA

SUMA applies shared nearest neighbour graph methods and a random forest model to optimize graph-based clustering and cell-type annotation of single-cell RNA sequencing (scRNA-Seq) data.


Key Features:

  • SNN graph representation: Represents scRNA-Seq data as a shared nearest neighbour (SNN) graph for graph-based clustering.
  • Graph-based clustering algorithms: Integrates three different graph-based clustering algorithms for comparative clustering analyses.
  • Random forest optimization: Uses a random forest model to predict the optimal number of neighbours required for effective clustering.
  • Parameter exploration: Evaluates clustering performance across a broad range of parameters including varying numbers of neighbours and variant genes.
  • Performance metrics: Reports model accuracy of 0.96 and an ROC Area Under the Curve (AUC) of 0.98.
  • Dataset scope: Developed and validated using publicly available scRNA-Seq datasets.
  • Feature importance: Identifies the number of cells as a critical feature influencing neighbour-count determination.
  • scRNA-Seq challenges addressed: Targets sparsity, noise, and uncertainty inherent to scRNA-Seq data to improve clustering and annotation.

Scientific Applications:

  • Clustering optimization: Optimize neighbour-parameter selection and clustering configurations for graph-based analyses of scRNA-Seq data.
  • Cell type annotation: Improve robustness of cell-type identification in single-cell transcriptomic studies.
  • Benchmarking and validation: Benchmark graph-based clustering methods across parameter ranges using accuracy and ROC AUC metrics.

Methodology:

Constructs a shared nearest neighbour (SNN) graph from scRNA-Seq data, applies three graph-based clustering algorithms, and uses a random forest model to predict the optimal number of neighbours; performance was evaluated across varying neighbour counts and variant genes using accuracy (0.96) and ROC AUC (0.98) on publicly available scRNA-Seq datasets, with feature importance identifying the number of cells as critical.

Topics

Details

Cost:
Free of charge
Tool Type:
web application
Operating Systems:
Mac, Linux, Windows
Programming Languages:
Python
Added:
6/18/2024
Last Updated:
11/24/2024

Operations

Publications

Koçhan N, Dayanc BE. Classification of colon cancer patients into consensus molecular subtypes using support vector machines. Turkish Journal of Biology. 2023;47(6):406-412. doi:10.55730/1300-0152.2675. PMID:38681777. PMCID:PMC11045205.