SRIQ
SRIQ performs unsupervised clustering of gene expression data to identify biologically meaningful tumor subtypes while avoiding pre-specification of the number of clusters and reducing subjectivity via feature pre-selection.
Key Features:
- Integration of Random Forest Concepts: Incorporates machine learning techniques from random forests to enhance clustering and pattern identification in complex gene expression datasets.
- Quality Threshold and k-Nearest Neighbor Clustering: Combines quality threshold methods with k-nearest neighbor clustering to identify core clusters of highly correlated observations and expand them into larger clusters.
- Automatic Evaluation of Stability: Autonomously evaluates clustering solution stability without requiring the user to pre-specify the number of cluster solutions to assess.
- Feature Pre-selection: Employs feature pre-selection to reduce subjectivity in input feature choice for clustering.
- Technical Reproducibility: Demonstrated reproducibility on 434 lung adenocarcinomas profiled by RNA sequencing.
- Comprehensive Pipeline Components: Includes data pre-processing, differential expression analysis, and pathway analysis as part of the analytical workflow.
Scientific Applications:
- Molecular subtype discovery: Identification of new molecular subtypes of tumors from gene expression data in cancer research.
- Refinement of transcriptional subtypes: Refinement of existing transcriptional subtypes defined by consensus clustering to provide more nuanced subclassification.
- Disease classification and clinical insight: Support for improved disease classification and for informing personalized treatment strategies based on transcriptional subtypes.
- Application to lung adenocarcinoma: Empirical application to RNA sequencing data from lung adenocarcinomas to define reproducible tumor subsets.
Methodology:
Uses feature pre-selection, integration of random forest concepts, quality threshold and k-nearest neighbor clustering, automatic evaluation of cluster stability, and includes data pre-processing, differential expression analysis, and pathway analysis.
Topics
Details
- License:
- GPL-3.0
- Cost:
- Free of charge
- Tool Type:
- command-line tool
- Operating Systems:
- Mac, Linux, Windows
- Programming Languages:
- Python, Java
- Added:
- 7/24/2022
- Last Updated:
- 11/24/2024
Operations
Publications
Karlström J, Aine M, Staaf J, Veerla S. SRIQ clustering: A fusion of Random Forest, QT clustering, and KNN concepts. Computational and Structural Biotechnology Journal. 2022;20:1567-1579. doi:10.1016/j.csbj.2022.03.036. PMID:35465158. PMCID:PMC9010551.