CLMB

CLMB applies a deep contrastive learning framework to metagenome binning to produce stable contig representations and improve robustness and accuracy of draft genome reconstruction from assembled contigs in metagenomic datasets.


Key Features:

  • Contrastive learning: Uses a deep contrastive learning framework to learn contig representations that are stable across perturbations.
  • Simulated noise injection: Introduces simulated noise into the training dataset to enforce representation consistency between noise-free and distorted data.
  • Implicit noise handling: Enhances robustness to real-world metagenomic noise without performing explicit denoising of input data.
  • Contig clustering: Clusters assembled contigs into draft genomes as the core binning operation.
  • Improved recovery: Recovers up to 17% more near-complete genomes across benchmarking datasets compared to the second-best method.
  • Enhanced bin refinement: Improves bin refinement, reconstructing an additional 8–22 high-quality genomes and 15–32 middle-quality genomes versus leading methods.
  • Refiner compatibility and comparison: Outperforms VAMB and Maxbin on benchmarking datasets, recovering an average of 15 more high-quality genomes.
  • Scalability on real data: Scales to real datasets, recovering 365 high-quality and middle-quality genomes, including 21 novel genomes, from a 110-sample mother-infant microbiome dataset.

Scientific Applications:

  • Genome reconstruction: Reconstruction of draft genomes from assembled contigs to recover uncultivated microbial populations and support downstream genome-resolved analyses.
  • Noise-robust binning: Metagenomic binning in datasets with technical and biological noise using representation stability enforced by contrastive learning.
  • Bin refinement and quality improvement: Enhancing bin refinement workflows to increase recovery of high-quality and middle-quality genomes.
  • Microbiome transmission studies: Application to mother-infant microbiome datasets for investigating transmission dynamics, as demonstrated on a 110-sample dataset recovering 365 genomes including 21 novel genomes.
  • Benchmarking and comparative evaluation: Comparative evaluation against state-of-the-art binners, providing quantitative improvements in near-complete and high-quality genome recovery.

Methodology:

Employs deep contrastive learning with simulated noise injection during training to learn stable, similar representations for noise-free and distorted contig data, followed by clustering assembled contigs into draft genomes and subsequent bin refinement.

Topics

Details

License:
MIT
Tool Type:
workflow
Programming Languages:
Python
Added:
3/28/2022
Last Updated:
3/28/2022

Operations

Publications

Zhang P, Jiang Z, Wang Y, Li Y. CLMB: deep contrastive learning for robust metagenomic binning. Unknown Journal. 2021. doi:10.1101/2021.11.15.468566.

Links