SRIQ

SRIQ performs unsupervised clustering of gene expression data to identify biologically meaningful tumor subtypes while avoiding pre-specification of the number of clusters and reducing subjectivity via feature pre-selection.


Key Features:

  • Integration of Random Forest Concepts: Incorporates machine learning techniques from random forests to enhance clustering and pattern identification in complex gene expression datasets.
  • Quality Threshold and k-Nearest Neighbor Clustering: Combines quality threshold methods with k-nearest neighbor clustering to identify core clusters of highly correlated observations and expand them into larger clusters.
  • Automatic Evaluation of Stability: Autonomously evaluates clustering solution stability without requiring the user to pre-specify the number of cluster solutions to assess.
  • Feature Pre-selection: Employs feature pre-selection to reduce subjectivity in input feature choice for clustering.
  • Technical Reproducibility: Demonstrated reproducibility on 434 lung adenocarcinomas profiled by RNA sequencing.
  • Comprehensive Pipeline Components: Includes data pre-processing, differential expression analysis, and pathway analysis as part of the analytical workflow.

Scientific Applications:

  • Molecular subtype discovery: Identification of new molecular subtypes of tumors from gene expression data in cancer research.
  • Refinement of transcriptional subtypes: Refinement of existing transcriptional subtypes defined by consensus clustering to provide more nuanced subclassification.
  • Disease classification and clinical insight: Support for improved disease classification and for informing personalized treatment strategies based on transcriptional subtypes.
  • Application to lung adenocarcinoma: Empirical application to RNA sequencing data from lung adenocarcinomas to define reproducible tumor subsets.

Methodology:

Uses feature pre-selection, integration of random forest concepts, quality threshold and k-nearest neighbor clustering, automatic evaluation of cluster stability, and includes data pre-processing, differential expression analysis, and pathway analysis.

Topics

Details

License:
GPL-3.0
Cost:
Free of charge
Tool Type:
command-line tool
Operating Systems:
Mac, Linux, Windows
Programming Languages:
Python, Java
Added:
7/24/2022
Last Updated:
11/24/2024

Operations

Publications

Karlström J, Aine M, Staaf J, Veerla S. SRIQ clustering: A fusion of Random Forest, QT clustering, and KNN concepts. Computational and Structural Biotechnology Journal. 2022;20:1567-1579. doi:10.1016/j.csbj.2022.03.036. PMID:35465158. PMCID:PMC9010551.