Matched Forest

Matched Forest identifies exposure variables and interaction effects in high-dimensional matched case-control studies by transforming matched data with counterfactuals and applying a modified variable importance score derived from supervised learning.


Key Features:

  • High-Dimensional Data Handling: Operates on datasets with large numbers of explanatory variables typical of modern biomedical studies.
  • Interaction Detection: Detects interaction effects among variables to capture complex joint associations.
  • Potential Outcome Model Basis: Leverages the potential outcome model framework to accommodate varying numbers of matching and exposure variables.
  • Counterfactual Data Transformation: Transforms matched case-control data by adding counterfactuals that preserve both case and control values for each instance.
  • Modified Variable Importance Score: Computes a modified variable importance score derived from supervised learning to prioritize significant predictors, including main effects and interactions.

Scientific Applications:

  • Variable selection in matched case-control studies: Identifies exposure variables associated with health outcomes within matched case-control designs.
  • Interaction discovery for exposure-outcome relationships: Detects interaction effects among exposures to support analyses of complex biological relationships and potential causal pathways.

Methodology:

The method transforms matched case-control data by adding counterfactuals within a potential outcome framework and then applies a supervised-learning-derived modified variable importance score to identify significant variables and interactions.

Topics

Details

Tool Type:
library
Programming Languages:
R
Added:
1/9/2020
Last Updated:
12/23/2020

Operations

Publications

Shomal Zadeh N, Lin S, Runger GC. Matched Forest: supervised learning for high-dimensional matched case–control studies. Bioinformatics. 2019;36(5):1570-1576. doi:10.1093/bioinformatics/btz785. PMID:31621830.