CANTARE
CANTARE constructs predictive regression models from network neighborhoods to integrate multi-omic data and produce interpretable outcome predictions.
Key Features:
- Pairwise regression modeling: Generates pairwise regression models across all analyte pairs from multiple omes.
- Network encoding: Encodes pairwise relationships into a network structure representing analyte interdependencies.
- Network-neighborhood predictors: Builds predictive logistic regression models centered on specific network neighborhoods.
- Interaction modeling: Supports inclusion of an interaction term for IBD in pairwise models to inform network construction.
- Quantitative directional effects: Reports quantitative and directional effect sizes for predictors to support hypothesis generation.
- Multi-omic support: Applies to datasets including gut microbiome, metabolomics, and microbial-derived enzymes.
- Model parsimony: Produces parsimonious models with predictor counts reported between 3 and 13 for top models.
- Performance metrics: Evaluates models using area under the curve (AUC) and statistical comparisons (including p-values).
- Method comparisons: Benchmarks against random forests and elastic-net penalized regressions and against correlation networks for performance assessment.
- Predictor prioritization and dynamic range: Demonstrates prioritization of predictors from multiple omes and a broader dynamic range in predicted probabilities (p = 1.35 × 10^-5; p = 0.005 reported for comparisons).
Scientific Applications:
- Disease outcome prediction: Building interpretable predictive models for inflammatory bowel disease (IBD) using multi-omic data.
- Multi-omic integration: Integrating gut microbiome, metabolomics, and microbial-derived enzyme data to identify cross-omic associations.
- Method benchmarking: Comparing network-based logistic regression models to random forests, elastic-net penalized regressions, and correlation networks using AUC and statistical tests.
- Hypothesis generation: Using quantitative directional effect sizes from models to generate testable biological hypotheses.
Methodology:
Generate pairwise regression models across all analyte pairs, encode significant pairwise relationships into a network, and construct predictive logistic regression models from network neighborhoods; optionally include an interaction term for IBD, evaluate model AUCs, and compare performance to random forests, elastic-net penalized regressions, and correlation-network-derived models using statistical tests (reported p-values).
Topics
Details
- Tool Type:
- workflow
- Programming Languages:
- R
- Added:
- 3/19/2021
- Last Updated:
- 4/22/2021
Operations
Publications
Siebert JC, Saint-Cyr M, Borengasser SJ, Wagner BD, Lozupone CA, Görg C. CANTARE: finding and visualizing network-based multi-omic predictive models. BMC Bioinformatics. 2021;22(1). doi:10.1186/s12859-021-04016-8. PMID:33607938. PMCID:PMC7896366.