Family Rank

Family Rank ranks and prioritizes features for predictive modeling by integrating graphical domain knowledge with empirical data to improve feature selection in high-dimensional, small-sample biological datasets.


Key Features:

  • Graphical Domain Knowledge Integration: Uses graphical models to incorporate domain-specific information into the feature ranking process.
  • Weighted Feature Scoring: Computes a weighted score for each feature that combines empirical data evidence and graphical knowledge.
  • Family-Based Scoring: Groups features into families and adjusts feature scores based on their presence and contributions across multiple families.
  • Overfitting Mitigation: Integrates domain knowledge and family-weighted scoring to reduce overfitting in high-dimensional, small-sample settings.

Scientific Applications:

  • Sample-size reduction for predictor detection: Demonstrated a 2- to 3-fold decrease in required sample size to detect true predictors compared to other state-of-the-art graph-based ranking algorithms.
  • Oncology biomarker and genetic-feature identification: Applicable to high-dimensional oncology datasets with limited samples for identifying biomarkers or genetic features associated with disease outcomes.

Methodology:

Uses graphical models to incorporate domain knowledge, computes per-feature weighted scores combining empirical data and graphical knowledge, sums family-weighted scores across all families a feature appears in, and selects features that maximize the resulting weighted score.

Topics

Details

License:
GPL-3.0
Tool Type:
library
Programming Languages:
R
Added:
9/8/2021
Last Updated:
9/13/2021

Operations

Publications

Saul M, Dinu V. Family Rank: a graphical domain knowledge informed feature ranking algorithm. Bioinformatics. 2021;37(20):3626-3631. doi:10.1093/bioinformatics/btab387. PMID:34009295.

Documentation