SUMA
SUMA applies shared nearest neighbour graph methods and a random forest model to optimize graph-based clustering and cell-type annotation of single-cell RNA sequencing (scRNA-Seq) data.
Key Features:
- SNN graph representation: Represents scRNA-Seq data as a shared nearest neighbour (SNN) graph for graph-based clustering.
- Graph-based clustering algorithms: Integrates three different graph-based clustering algorithms for comparative clustering analyses.
- Random forest optimization: Uses a random forest model to predict the optimal number of neighbours required for effective clustering.
- Parameter exploration: Evaluates clustering performance across a broad range of parameters including varying numbers of neighbours and variant genes.
- Performance metrics: Reports model accuracy of 0.96 and an ROC Area Under the Curve (AUC) of 0.98.
- Dataset scope: Developed and validated using publicly available scRNA-Seq datasets.
- Feature importance: Identifies the number of cells as a critical feature influencing neighbour-count determination.
- scRNA-Seq challenges addressed: Targets sparsity, noise, and uncertainty inherent to scRNA-Seq data to improve clustering and annotation.
Scientific Applications:
- Clustering optimization: Optimize neighbour-parameter selection and clustering configurations for graph-based analyses of scRNA-Seq data.
- Cell type annotation: Improve robustness of cell-type identification in single-cell transcriptomic studies.
- Benchmarking and validation: Benchmark graph-based clustering methods across parameter ranges using accuracy and ROC AUC metrics.
Methodology:
Constructs a shared nearest neighbour (SNN) graph from scRNA-Seq data, applies three graph-based clustering algorithms, and uses a random forest model to predict the optimal number of neighbours; performance was evaluated across varying neighbour counts and variant genes using accuracy (0.96) and ROC AUC (0.98) on publicly available scRNA-Seq datasets, with feature importance identifying the number of cells as critical.
Topics
Details
- Cost:
- Free of charge
- Tool Type:
- web application
- Operating Systems:
- Mac, Linux, Windows
- Programming Languages:
- Python
- Added:
- 6/18/2024
- Last Updated:
- 11/24/2024
Operations
Publications
Koçhan N, Dayanc BE. Classification of colon cancer patients into consensus molecular subtypes using support vector machines. Turkish Journal of Biology. 2023;47(6):406-412. doi:10.55730/1300-0152.2675. PMID:38681777. PMCID:PMC11045205.