CLMB
CLMB applies a deep contrastive learning framework to metagenome binning to produce stable contig representations and improve robustness and accuracy of draft genome reconstruction from assembled contigs in metagenomic datasets.
Key Features:
- Contrastive learning: Uses a deep contrastive learning framework to learn contig representations that are stable across perturbations.
- Simulated noise injection: Introduces simulated noise into the training dataset to enforce representation consistency between noise-free and distorted data.
- Implicit noise handling: Enhances robustness to real-world metagenomic noise without performing explicit denoising of input data.
- Contig clustering: Clusters assembled contigs into draft genomes as the core binning operation.
- Improved recovery: Recovers up to 17% more near-complete genomes across benchmarking datasets compared to the second-best method.
- Enhanced bin refinement: Improves bin refinement, reconstructing an additional 8–22 high-quality genomes and 15–32 middle-quality genomes versus leading methods.
- Refiner compatibility and comparison: Outperforms VAMB and Maxbin on benchmarking datasets, recovering an average of 15 more high-quality genomes.
- Scalability on real data: Scales to real datasets, recovering 365 high-quality and middle-quality genomes, including 21 novel genomes, from a 110-sample mother-infant microbiome dataset.
Scientific Applications:
- Genome reconstruction: Reconstruction of draft genomes from assembled contigs to recover uncultivated microbial populations and support downstream genome-resolved analyses.
- Noise-robust binning: Metagenomic binning in datasets with technical and biological noise using representation stability enforced by contrastive learning.
- Bin refinement and quality improvement: Enhancing bin refinement workflows to increase recovery of high-quality and middle-quality genomes.
- Microbiome transmission studies: Application to mother-infant microbiome datasets for investigating transmission dynamics, as demonstrated on a 110-sample dataset recovering 365 genomes including 21 novel genomes.
- Benchmarking and comparative evaluation: Comparative evaluation against state-of-the-art binners, providing quantitative improvements in near-complete and high-quality genome recovery.
Methodology:
Employs deep contrastive learning with simulated noise injection during training to learn stable, similar representations for noise-free and distorted contig data, followed by clustering assembled contigs into draft genomes and subsequent bin refinement.
Topics
Details
- License:
- MIT
- Tool Type:
- workflow
- Programming Languages:
- Python
- Added:
- 3/28/2022
- Last Updated:
- 3/28/2022
Operations
Publications
Zhang P, Jiang Z, Wang Y, Li Y. CLMB: deep contrastive learning for robust metagenomic binning. Unknown Journal. 2021. doi:10.1101/2021.11.15.468566.
Links
Issue tracker
https://github.com/zpf0117b/CLMB/issues