PermFIT

PermFIT implements permutation-based feature importance testing within machine learning models to identify and interpret biomarkers across genetic, genomic, and proteomic data and is provided as part of the deepTL R package.


Key Features:

  • Permutation-based testing: Uses permutation-based feature importance to estimate and test feature effects.
  • Model compatibility: Operates with deep neural networks, random forests, and support vector machines.
  • No model refitting required: Provides interpretable feature importance estimates without necessitating model refitting.
  • Computational efficiency: Designed to be computationally efficient for application to large datasets while supporting robust statistical inference.
  • Statistical validity: Demonstrated through numerical studies to yield valid statistical inferences and to enhance prediction accuracy.
  • Implementation: Implemented within the deepTL R package.
  • Empirical application: Applied to datasets such as the Cancer Genome Atlas kidney tumor data and the HITChip atlas data to identify significant biomarkers.

Scientific Applications:

  • Biomarker discovery: Identification and interpretation of important biomarkers in complex human diseases at genetic, genomic, and proteomic levels.
  • Model interpretation: Facilitating transparent interpretation of sophisticated machine learning algorithms.
  • Prediction improvement: Improving prediction accuracy of machine learning models across diverse scenarios.
  • Translational hypothesis generation: Generating hypotheses relevant to prevention, diagnosis, and treatment of complex diseases.
  • Real-data analysis: Analysis of Cancer Genome Atlas kidney tumor data and HITChip atlas data for biomarker identification.

Methodology:

PermFIT applies a permutation-based approach to estimate and test feature importance across machine learning frameworks (deep neural networks, random forests, support vector machines) and provides interpretable feature importance measures without model refitting; its performance has been evaluated via numerical studies and applications to real datasets.

Topics

Details

Tool Type:
library
Programming Languages:
R, C++
Added:
11/1/2021
Last Updated:
11/1/2021

Operations

Publications

Mi X, Zou B, Zou F, Hu J. Permutation-based identification of important biomarkers for complex diseases via machine learning models. Nature Communications. 2021;12(1). doi:10.1038/s41467-021-22756-2. PMID:34021151. PMCID:PMC8140109.

Downloads

Links