PermFIT
PermFIT implements permutation-based feature importance testing within machine learning models to identify and interpret biomarkers across genetic, genomic, and proteomic data and is provided as part of the deepTL R package.
Key Features:
- Permutation-based testing: Uses permutation-based feature importance to estimate and test feature effects.
- Model compatibility: Operates with deep neural networks, random forests, and support vector machines.
- No model refitting required: Provides interpretable feature importance estimates without necessitating model refitting.
- Computational efficiency: Designed to be computationally efficient for application to large datasets while supporting robust statistical inference.
- Statistical validity: Demonstrated through numerical studies to yield valid statistical inferences and to enhance prediction accuracy.
- Implementation: Implemented within the deepTL R package.
- Empirical application: Applied to datasets such as the Cancer Genome Atlas kidney tumor data and the HITChip atlas data to identify significant biomarkers.
Scientific Applications:
- Biomarker discovery: Identification and interpretation of important biomarkers in complex human diseases at genetic, genomic, and proteomic levels.
- Model interpretation: Facilitating transparent interpretation of sophisticated machine learning algorithms.
- Prediction improvement: Improving prediction accuracy of machine learning models across diverse scenarios.
- Translational hypothesis generation: Generating hypotheses relevant to prevention, diagnosis, and treatment of complex diseases.
- Real-data analysis: Analysis of Cancer Genome Atlas kidney tumor data and HITChip atlas data for biomarker identification.
Methodology:
PermFIT applies a permutation-based approach to estimate and test feature importance across machine learning frameworks (deep neural networks, random forests, support vector machines) and provides interpretable feature importance measures without model refitting; its performance has been evaluated via numerical studies and applications to real datasets.
Topics
Details
- Tool Type:
- library
- Programming Languages:
- R, C++
- Added:
- 11/1/2021
- Last Updated:
- 11/1/2021
Operations
Publications
Mi X, Zou B, Zou F, Hu J. Permutation-based identification of important biomarkers for complex diseases via machine learning models. Nature Communications. 2021;12(1). doi:10.1038/s41467-021-22756-2. PMID:34021151. PMCID:PMC8140109.