MetaCluster
MetaCluster annotates metagenomic reads using assembly-assisted binning to create virtual contigs that improve taxonomic annotation of short-read (~100 bp) metagenomic data.
Key Features:
- Assembly-assisted binning-based annotation: Groups reads or contigs into bins and uses aggregated bin information to perform taxonomic annotation rather than aligning individual short reads alone.
- Virtual contigs: Represents sets of binned reads as "virtual contigs" that can reach up to 10 Kb in length to enable more effective alignments.
- Cluster representation: Aggregates multiple virtual contigs within a cluster to represent genomic content up to 1 Mb per cluster.
- Similarity-based clustering: Clusters reads or contigs based on their similarity with the assumption that reads in a cluster originate from the same taxon.
- Reduced alignment volume: Decreases the number of sequences requiring alignment by aligning longer virtual contigs instead of many short reads.
- Improved annotation accuracy: Leverages extended sequence length and comprehensive cluster representation to increase the number and accuracy of annotated reads.
- Enhanced binning precision with reference databases: Refines binning results using reference taxonomic databases to improve taxonomic resolution.
- Comparative performance: Reports higher read annotation counts, greater accuracy, and improved efficiency relative to MEGAN4 and MetaCluster 5.0.
Scientific Applications:
- Taxonomic annotation of metagenomes: Assigns taxonomy to short-read metagenomic sequences by aligning virtual contigs to taxonomic databases.
- Binning and genomic representation: Produces bins and cluster-level representations of genomic content for downstream analyses.
- Mitigation of short-read limitations: Addresses challenges of ~100 bp reads and immature assembly by aggregating reads into longer virtual contigs.
- Method comparison and benchmarking: Enables comparative assessment of annotation and binning performance against tools such as MEGAN4 and MetaCluster 5.0.
Methodology:
Reads or contigs are grouped into clusters based on similarity; binned reads are aggregated into virtual contigs up to 10 Kb and clusters up to 1 Mb; these virtual contigs/clusters are aligned to reference taxonomic databases and binning is refined using those reference taxonomic databases.
Topics
Details
- Tool Type:
- command-line tool
- Operating Systems:
- Linux
- Programming Languages:
- C++
- Added:
- 8/3/2017
- Last Updated:
- 11/25/2024
Operations
Publications
Wang Y, Leung HCM, Yiu SM, Chin FYL. MetaCluster-TA: taxonomic annotation for metagenomic data based on assembly-assisted binning. BMC Genomics. 2014;15(S1). doi:10.1186/1471-2164-15-s1-s12. PMID:24564377. PMCID:PMC4046714.