Peculiar Genes Selection (PGS)

Peculiar Genes Selection (PGS) identifies class-specific biomarker genes from high-dimensional genomics and transcriptomics data to improve classification performance, particularly in imbalanced datasets.


Key Features:

  • Three-Step Feature Selection Methodology: Performs differential expression detection, filters features by discriminative power, and identifies class-specific features that form biomarker signatures.
  • Handling Imbalanced Data: Targets datasets with imbalanced class sizes and has been demonstrated to improve classification performance using Support Vector Machines (SVM), achieving 82% accuracy versus ~73% for comparative methods.
  • Performance on Balanced Datasets: Shows comparable performance to existing methods when applied to perfectly balanced datasets.
  • Biological Relevance Validation: Validates selected biomarker signatures using Gene Ontology enrichment analyses to confirm biological relevance.

Scientific Applications:

  • Genomics and Transcriptomics: Identification of class-specific biomarkers in high-dimensional genomics and transcriptomics datasets.
  • Disease Classification: Selection of discriminative gene signatures to support disease class prediction.
  • Personalized Medicine: Enabling more accurate patient stratification based on selected biomarker signatures.
  • Biomarker Discovery: Facilitating discovery and validation of biomarkers for downstream biological interpretation and functional analysis.

Methodology:

Identifies differentially expressed genes, filters features by discriminative power, detects class-specific features, performs classification with Support Vector Machines (SVM), and evaluates biological relevance via Gene Ontology enrichment.

Topics

Details

Tool Type:
library
Operating Systems:
Linux, Windows, Mac
Programming Languages:
R
Added:
7/7/2018
Last Updated:
11/25/2024

Operations

Publications

Martina F, Beccuti M, Balbo G, Cordero F. Peculiar Genes Selection: A new features selection method to improve classification performances in imbalanced data sets. PLOS ONE. 2017;12(8):e0177475. doi:10.1371/journal.pone.0177475. PMID:28806759. PMCID:PMC5555681.

Documentation