CANTARE

CANTARE constructs predictive regression models from network neighborhoods to integrate multi-omic data and produce interpretable outcome predictions.


Key Features:

  • Pairwise regression modeling: Generates pairwise regression models across all analyte pairs from multiple omes.
  • Network encoding: Encodes pairwise relationships into a network structure representing analyte interdependencies.
  • Network-neighborhood predictors: Builds predictive logistic regression models centered on specific network neighborhoods.
  • Interaction modeling: Supports inclusion of an interaction term for IBD in pairwise models to inform network construction.
  • Quantitative directional effects: Reports quantitative and directional effect sizes for predictors to support hypothesis generation.
  • Multi-omic support: Applies to datasets including gut microbiome, metabolomics, and microbial-derived enzymes.
  • Model parsimony: Produces parsimonious models with predictor counts reported between 3 and 13 for top models.
  • Performance metrics: Evaluates models using area under the curve (AUC) and statistical comparisons (including p-values).
  • Method comparisons: Benchmarks against random forests and elastic-net penalized regressions and against correlation networks for performance assessment.
  • Predictor prioritization and dynamic range: Demonstrates prioritization of predictors from multiple omes and a broader dynamic range in predicted probabilities (p = 1.35 × 10^-5; p = 0.005 reported for comparisons).

Scientific Applications:

  • Disease outcome prediction: Building interpretable predictive models for inflammatory bowel disease (IBD) using multi-omic data.
  • Multi-omic integration: Integrating gut microbiome, metabolomics, and microbial-derived enzyme data to identify cross-omic associations.
  • Method benchmarking: Comparing network-based logistic regression models to random forests, elastic-net penalized regressions, and correlation networks using AUC and statistical tests.
  • Hypothesis generation: Using quantitative directional effect sizes from models to generate testable biological hypotheses.

Methodology:

Generate pairwise regression models across all analyte pairs, encode significant pairwise relationships into a network, and construct predictive logistic regression models from network neighborhoods; optionally include an interaction term for IBD, evaluate model AUCs, and compare performance to random forests, elastic-net penalized regressions, and correlation-network-derived models using statistical tests (reported p-values).

Topics

Details

Tool Type:
workflow
Programming Languages:
R
Added:
3/19/2021
Last Updated:
4/22/2021

Operations

Publications

Siebert JC, Saint-Cyr M, Borengasser SJ, Wagner BD, Lozupone CA, Görg C. CANTARE: finding and visualizing network-based multi-omic predictive models. BMC Bioinformatics. 2021;22(1). doi:10.1186/s12859-021-04016-8. PMID:33607938. PMCID:PMC7896366.