MetaCluster

MetaCluster annotates metagenomic reads using assembly-assisted binning to create virtual contigs that improve taxonomic annotation of short-read (~100 bp) metagenomic data.


Key Features:

  • Assembly-assisted binning-based annotation: Groups reads or contigs into bins and uses aggregated bin information to perform taxonomic annotation rather than aligning individual short reads alone.
  • Virtual contigs: Represents sets of binned reads as "virtual contigs" that can reach up to 10 Kb in length to enable more effective alignments.
  • Cluster representation: Aggregates multiple virtual contigs within a cluster to represent genomic content up to 1 Mb per cluster.
  • Similarity-based clustering: Clusters reads or contigs based on their similarity with the assumption that reads in a cluster originate from the same taxon.
  • Reduced alignment volume: Decreases the number of sequences requiring alignment by aligning longer virtual contigs instead of many short reads.
  • Improved annotation accuracy: Leverages extended sequence length and comprehensive cluster representation to increase the number and accuracy of annotated reads.
  • Enhanced binning precision with reference databases: Refines binning results using reference taxonomic databases to improve taxonomic resolution.
  • Comparative performance: Reports higher read annotation counts, greater accuracy, and improved efficiency relative to MEGAN4 and MetaCluster 5.0.

Scientific Applications:

  • Taxonomic annotation of metagenomes: Assigns taxonomy to short-read metagenomic sequences by aligning virtual contigs to taxonomic databases.
  • Binning and genomic representation: Produces bins and cluster-level representations of genomic content for downstream analyses.
  • Mitigation of short-read limitations: Addresses challenges of ~100 bp reads and immature assembly by aggregating reads into longer virtual contigs.
  • Method comparison and benchmarking: Enables comparative assessment of annotation and binning performance against tools such as MEGAN4 and MetaCluster 5.0.

Methodology:

Reads or contigs are grouped into clusters based on similarity; binned reads are aggregated into virtual contigs up to 10 Kb and clusters up to 1 Mb; these virtual contigs/clusters are aligned to reference taxonomic databases and binning is refined using those reference taxonomic databases.

Topics

Details

Tool Type:
command-line tool
Operating Systems:
Linux
Programming Languages:
C++
Added:
8/3/2017
Last Updated:
11/25/2024

Operations

Publications

Wang Y, Leung HCM, Yiu SM, Chin FYL. MetaCluster-TA: taxonomic annotation for metagenomic data based on assembly-assisted binning. BMC Genomics. 2014;15(S1). doi:10.1186/1471-2164-15-s1-s12. PMID:24564377. PMCID:PMC4046714.

Documentation

Links