SamCluster
SamCluster identifies sample classes from gene expression profiles by integrating coefficient of variation (CV)–based and t-test feature selection with hierarchical clustering.
Key Features:
- Feature Selection Algorithms: Selects genes based on coefficient of variation (CV) and significance from t-tests to retain genes relevant for clustering.
- Hierarchical Clustering Integration: Integrates feature selection with hierarchical clustering, using CV thresholds to facilitate initial separation of samples into two groups.
- Iterative Refinement: Iteratively identifies significantly differentially expressed genes using t-tests with p-values ≤ 0.01, 0.05, or 0.1 until stable sample classes emerge.
- Consensus Class Construction: Constructs consensus sample classes from putative classes derived at different CV thresholds and selects optimal class pairs by minimizing distance between consensus and putative classes.
Scientific Applications:
- Tissue sample classification: Predicts and refines sample class labels from gene expression datasets for tissue-level analyses.
- Oncology and immunology: Supports studies in oncology and immunology by identifying gene-expression–defined sample classes relevant to those fields.
- Validation datasets: Has been applied to COLON, LEUKEMIA72, LEUKEMIA38, and OVARIAN datasets with reported misclassification counts of 5, 1, 0, and 0 samples respectively.
Methodology:
Select genes whose coefficient of variation (CV) exceeds thresholds; apply hierarchical clustering to split samples (initially into two groups); iteratively identify differentially expressed genes using t-tests (p ≤ 0.01, 0.05, or 0.1) until classes stabilize; construct consensus classes from putative classes across CV thresholds and choose optimal class pairs by minimizing distance between consensus and putative classes.
Topics
Details
- Tool Type:
- plugin
- Operating Systems:
- Linux, Windows, Mac
- Added:
- 8/3/2017
- Last Updated:
- 11/25/2024
Operations
Publications
Li W, Fan M, Xiong M. SamCluster: an integrated scheme for automatic discovery of sample classes using gene expression profile. Bioinformatics. 2003;19(7):811-817. doi:10.1093/bioinformatics/btg095. PMID:12724290.