TADA

TADA augments microbiome training datasets with taxonomy-aware synthetic samples to improve phenotype classification from high-dimensional, low-sample-size microbial profiles.


Key Features:

  • Phylogenetic Augmentation: Leverages phylogenetic relationships among microbial taxa to guide generation of biologically relevant synthetic samples.
  • Statistical Generative Model: Uses a statistical generative model to create realistic microbiome profiles for augmentation.
  • Synthetic Sample Generation: Generates synthetic samples that preserve evolutionary connections and biological integrity.
  • Addressing Data Imbalance: Augments under-represented classes to mitigate class imbalance and bias in microbiome datasets.
  • Improved Machine Learning Robustness: Enhances robustness and accuracy of machine learning models applied to phenotype classification on high-dimensional, low-sample-size data.

Scientific Applications:

  • Phenotype Classification: Improves predictive performance of models associating microbial compositions with specific host traits.
  • Class Imbalance Scenarios: Applied where class imbalance and insufficient signal hinder traditional machine learning in microbiome studies.

Methodology:

Leverages phylogenetic relationships to generate synthetic microbiome samples via a statistical generative model and augments under-represented classes to balance training data.

Topics

Details

Programming Languages:
R, Shell, Python
Added:
11/14/2019
Last Updated:
12/27/2020

Operations

Publications

Sayyari E, Kawas B, Mirarab S. TADA: phylogenetic augmentation of microbiome samples enhances phenotype classification. Bioinformatics. 2019;35(14):i31-i40. doi:10.1093/bioinformatics/btz394. PMID:31510701. PMCID:PMC6612822.

PMID: 31510701
PMCID: PMC6612822
Funding: - National Science Foundation: III-1845967, IIS-1565862