Family Rank
Family Rank ranks and prioritizes features for predictive modeling by integrating graphical domain knowledge with empirical data to improve feature selection in high-dimensional, small-sample biological datasets.
Key Features:
- Graphical Domain Knowledge Integration: Uses graphical models to incorporate domain-specific information into the feature ranking process.
- Weighted Feature Scoring: Computes a weighted score for each feature that combines empirical data evidence and graphical knowledge.
- Family-Based Scoring: Groups features into families and adjusts feature scores based on their presence and contributions across multiple families.
- Overfitting Mitigation: Integrates domain knowledge and family-weighted scoring to reduce overfitting in high-dimensional, small-sample settings.
Scientific Applications:
- Sample-size reduction for predictor detection: Demonstrated a 2- to 3-fold decrease in required sample size to detect true predictors compared to other state-of-the-art graph-based ranking algorithms.
- Oncology biomarker and genetic-feature identification: Applicable to high-dimensional oncology datasets with limited samples for identifying biomarkers or genetic features associated with disease outcomes.
Methodology:
Uses graphical models to incorporate domain knowledge, computes per-feature weighted scores combining empirical data and graphical knowledge, sums family-weighted scores across all families a feature appears in, and selects features that maximize the resulting weighted score.
Topics
Details
- License:
- GPL-3.0
- Tool Type:
- library
- Programming Languages:
- R
- Added:
- 9/8/2021
- Last Updated:
- 9/13/2021
Operations
Publications
Saul M, Dinu V. Family Rank: a graphical domain knowledge informed feature ranking algorithm. Bioinformatics. 2021;37(20):3626-3631. doi:10.1093/bioinformatics/btab387. PMID:34009295.
PMID: 34009295