CoDaCoRe

CoDaCoRe identifies sparse, interpretable, and predictive log-ratio biomarkers from high-throughput sequencing (HTS) compositional data to analyze relationships among compositional variables and predict biological outcomes.


Key Features:

  • Sparse log-ratio biomarker selection: Produces sparse, interpretable log-ratio balances as candidate biomarkers.
  • Continuous relaxation of combinatorial optimization: Converts the discrete log-ratio selection problem into a continuous relaxation for optimization.
  • Gradient descent optimization: Optimizes the relaxed objective using gradient descent.
  • Deep learning integration: Leverages deep learning techniques to enable optimization and model learning.
  • Scalability to high-dimensional HTS/CoDa data: Designed to scale to high-dimensional compositional datasets such as metagenomic data.
  • State-of-the-art predictive accuracy and sparsity: Maintains competitive predictive performance while enforcing sparsity in selected log-ratios.
  • Computational speed: Operates several orders of magnitude faster than existing sparse log-ratio selection methods.
  • Benchmark validation: Validated across microbiome, metabolite, and microRNA benchmark datasets.

Scientific Applications:

  • Microbiome biomarker discovery: Identification of log-ratio biomarkers from microbiome HTS datasets.
  • Metabolomics biomarker discovery: Discovery of predictive log-ratio features in metabolite datasets.
  • microRNA biomarker discovery: Selection of interpretable log-ratio biomarkers from microRNA data.
  • Compositional data analysis and predictive modeling: Analysis of compositional (CoDa) HTS data for interpretable predictive models in high-dimensional settings.

Methodology:

Applies a continuous relaxation of the combinatorial log-ratio selection problem and optimizes the resulting objective using gradient descent within a deep learning framework.

Topics

Details

License:
Other
Cost:
Free of charge
Tool Type:
library
Operating Systems:
Mac, Linux, Windows
Programming Languages:
R
Added:
2/14/2022
Last Updated:
11/24/2024

Operations

Publications

Gordon-Rodriguez E, Quinn TP, Cunningham JP. Learning sparse log-ratios for high-throughput sequencing data. Bioinformatics. 2021;38(1):157-163. doi:10.1093/bioinformatics/btab645. PMID:34498030. PMCID:PMC8696089.

PMID: 34498030
PMCID: PMC8696089
Funding: - Simons Foundation: 542963 - NSF: DBI-1707398

Documentation

Links