SamCluster

SamCluster identifies sample classes from gene expression profiles by integrating coefficient of variation (CV)–based and t-test feature selection with hierarchical clustering.


Key Features:

  • Feature Selection Algorithms: Selects genes based on coefficient of variation (CV) and significance from t-tests to retain genes relevant for clustering.
  • Hierarchical Clustering Integration: Integrates feature selection with hierarchical clustering, using CV thresholds to facilitate initial separation of samples into two groups.
  • Iterative Refinement: Iteratively identifies significantly differentially expressed genes using t-tests with p-values ≤ 0.01, 0.05, or 0.1 until stable sample classes emerge.
  • Consensus Class Construction: Constructs consensus sample classes from putative classes derived at different CV thresholds and selects optimal class pairs by minimizing distance between consensus and putative classes.

Scientific Applications:

  • Tissue sample classification: Predicts and refines sample class labels from gene expression datasets for tissue-level analyses.
  • Oncology and immunology: Supports studies in oncology and immunology by identifying gene-expression–defined sample classes relevant to those fields.
  • Validation datasets: Has been applied to COLON, LEUKEMIA72, LEUKEMIA38, and OVARIAN datasets with reported misclassification counts of 5, 1, 0, and 0 samples respectively.

Methodology:

Select genes whose coefficient of variation (CV) exceeds thresholds; apply hierarchical clustering to split samples (initially into two groups); iteratively identify differentially expressed genes using t-tests (p ≤ 0.01, 0.05, or 0.1) until classes stabilize; construct consensus classes from putative classes across CV thresholds and choose optimal class pairs by minimizing distance between consensus and putative classes.

Topics

Details

Tool Type:
plugin
Operating Systems:
Linux, Windows, Mac
Added:
8/3/2017
Last Updated:
11/25/2024

Operations

Publications

Li W, Fan M, Xiong M. SamCluster: an integrated scheme for automatic discovery of sample classes using gene expression profile. Bioinformatics. 2003;19(7):811-817. doi:10.1093/bioinformatics/btg095. PMID:12724290.

Documentation

Links