MaAsLin2

MaAsLin2 identifies multivariable associations between phenotypic traits, environmental exposures, and microbial meta-omics features to enable rigorous analysis of noisy, sparse (zero-inflated), high-dimensional count and compositional microbiome data.


Key Features:

  • Statistical models: Employs generalized linear models (GLMs) and mixed models to model associations.
  • Data types: Analyzes counts and relative abundances (compositional measurements) from meta-omics datasets.
  • Zero-inflation and sparsity: Explicitly addresses noisy, sparse, and zero-inflated data characteristics.
  • Non-normality and high-dimensionality: Designed for extreme non-normality and high-dimensional feature spaces.
  • Study designs: Supports cross-sectional and longitudinal (repeated-measures) epidemiological study designs.
  • Covariate adjustment: Accommodates multiple covariates and repeated measures within models.
  • Statistical power: Maintains power in analyses with repeated measurements and multiple covariates.
  • False discovery control: Controls false discovery rates to provide robust results.
  • Evaluation: Optimized and assessed through simulation studies tailored to meta-omics association challenges.
  • Data processing: Includes approaches for data normalization, transformation, and exploration.
  • Demonstrated application: Applied to the Integrative Human Microbiome Project (HMP2) microbial multi-omics dataset to reproduce findings and identify novel insights into inflammatory bowel diseases across time points and omics profiles.

Scientific Applications:

  • Multivariable association discovery: Identifies associations between microbial features and phenotypic traits or environmental exposures in meta-omics studies.
  • Observational microbiome studies: Supports population-scale cross-sectional and longitudinal observational study analyses.
  • Multi-omics integration for disease: Enables integrated analysis of microbial multi-omics datasets such as HMP2 to study inflammatory bowel disease dynamics across time and omics layers.

Methodology:

MaAsLin2 employs generalized linear models (GLMs) and mixed models to analyze counts and relative abundances while accommodating covariates and repeated measures, performs data normalization, transformation, and exploration, uses simulation studies for evaluation, and controls false discovery rates.

Topics

Details

License:
Other
Programming Languages:
R
Added:
3/13/2024
Last Updated:
11/6/2024

Operations

Publications

Mallick H, Rahnavard A, McIver LJ, Ma S, Zhang Y, Nguyen LH, Tickle TL, Weingart G, Ren B, Schwager EH, Chatterjee S, Thompson KN, Wilkinson JE, Subramanian A, Lu Y, Waldron L, Paulson JN, Franzosa EA, Bravo HC, Huttenhower C. Multivariable association discovery in population-scale meta-omics studies. PLOS Computational Biology. 2021;17(11):e1009442. doi:10.1371/journal.pcbi.1009442. PMID:34784344. PMCID:PMC8714082.

Funding: - US National Science Foundation, Division of Environmental Biology: DEB 2028280 - national institute of allergy and infectious diseases: U19AI110820 - national human genome research institute: R01HG005220 - national institute of diabetes and digestive and kidney diseases: R24DK110499, U54DK102557

Documentation