Peculiar Genes Selection (PGS)
Peculiar Genes Selection (PGS) identifies class-specific biomarker genes from high-dimensional genomics and transcriptomics data to improve classification performance, particularly in imbalanced datasets.
Key Features:
- Three-Step Feature Selection Methodology: Performs differential expression detection, filters features by discriminative power, and identifies class-specific features that form biomarker signatures.
- Handling Imbalanced Data: Targets datasets with imbalanced class sizes and has been demonstrated to improve classification performance using Support Vector Machines (SVM), achieving 82% accuracy versus ~73% for comparative methods.
- Performance on Balanced Datasets: Shows comparable performance to existing methods when applied to perfectly balanced datasets.
- Biological Relevance Validation: Validates selected biomarker signatures using Gene Ontology enrichment analyses to confirm biological relevance.
Scientific Applications:
- Genomics and Transcriptomics: Identification of class-specific biomarkers in high-dimensional genomics and transcriptomics datasets.
- Disease Classification: Selection of discriminative gene signatures to support disease class prediction.
- Personalized Medicine: Enabling more accurate patient stratification based on selected biomarker signatures.
- Biomarker Discovery: Facilitating discovery and validation of biomarkers for downstream biological interpretation and functional analysis.
Methodology:
Identifies differentially expressed genes, filters features by discriminative power, detects class-specific features, performs classification with Support Vector Machines (SVM), and evaluates biological relevance via Gene Ontology enrichment.
Topics
Details
- Tool Type:
- library
- Operating Systems:
- Linux, Windows, Mac
- Programming Languages:
- R
- Added:
- 7/7/2018
- Last Updated:
- 11/25/2024
Operations
Publications
Martina F, Beccuti M, Balbo G, Cordero F. Peculiar Genes Selection: A new features selection method to improve classification performances in imbalanced data sets. PLOS ONE. 2017;12(8):e0177475. doi:10.1371/journal.pone.0177475. PMID:28806759. PMCID:PMC5555681.