MOCAT

MOCAT processes shotgun metagenomics Illumina sequencing data to perform quality control, assembly, gene prediction, and taxonomic and functional profiling of microbial communities.


Key Features:

  • Modular architecture: Modular architecture permitting customization and interchange of processing components.
  • Raw data handling: Read preprocessing of raw fastQ files including quality control to filter low-quality reads and contaminants.
  • Assembly: Assembly of metagenomic reads into contigs using assembly programs optimized for high-throughput sequencing.
  • Gene prediction: Prediction of protein-coding genes on assembled contigs using advanced gene prediction algorithms.
  • Taxonomic and functional profiling: Taxonomic and functional profiling by mapping reads against reference databases, enabling identification of phylogenetic marker genes, extraction or removal of specific sequences, and abundance calculations.
  • MOCAT2 reference catalogs: Automated generation and annotation of non-redundant reference catalogs built from pre-computed assignments across 18 databases covering diverse functional categories (MOCAT2).
  • HPC integration: Integration with LSF, PBS, and SGE queuing systems for high-performance computing on UNIX-based platforms.
  • Implementation: Implemented in Perl 5 and Python 2.7.
  • Output formats: Produces multi-sheet Excel documents and queryable SQL databases summarizing statistics from each processing step.
  • License: Distributed under the GPL3 open-source license.
  • Benchmarking: Benchmarked on artificial, real, and simulated metagenomes with reported improvements in selected quality metrics.

Scientific Applications:

  • Shotgun metagenomics: Processing and analysis of shotgun metagenomics datasets generated with Illumina sequencing to profile microbial community composition and function.
  • Gene catalog construction: Assembly of reads and prediction of protein-coding genes to construct gene catalogs and reference databases.
  • Functional annotation: Functional annotation and characterization of metagenomes using non-redundant catalogs annotated across 18 functional databases.
  • Taxonomic profiling: Generation of taxonomic profiles via identification of phylogenetic marker genes and read mapping.
  • Abundance estimation and sequence filtering: Abundance estimation of taxa and genes and removal or extraction of specific sequences.
  • Method benchmarking: Benchmarking and validation of metagenomic workflows using artificial, real, and simulated metagenomes.

Methodology:

Computational steps explicitly include raw fastQ quality control and contaminant filtering, assembly of reads into contigs with assembly programs, protein-coding gene prediction on contigs, read mapping to reference databases for taxonomic and functional profiling (including abundance calculations and sequence extraction/removal), automated creation and annotation of non-redundant reference catalogs from pre-computed assignments across 18 databases (MOCAT2), and execution via LSF, PBS, or SGE on UNIX; implemented in Perl 5 and Python 2.7 with outputs as multi-sheet Excel and SQL databases.

Topics

Collections

Details

License:
GPL-3.0
Maturity:
Mature
Cost:
Free of charge
Tool Type:
command-line tool
Operating Systems:
Linux, Mac
Programming Languages:
Perl, Python
Added:
4/5/2017
Last Updated:
11/24/2024

Operations

Data Inputs & Outputs

Community profiling

Publications

Kultima JR, Coelho LP, Forslund K, Huerta-Cepas J, Li SS, Driessen M, Voigt AY, Zeller G, Sunagawa S, Bork P. MOCAT2: a metagenomic assembly, annotation and profiling framework. Bioinformatics. 2016;32(16):2520-2523. doi:10.1093/bioinformatics/btw183. PMID:27153620. PMCID:PMC4978931.

Kultima JR, Sunagawa S, Li J, Chen W, Chen H, Mende DR, Arumugam M, Pan Q, Liu B, Qin J, Wang J, Bork P. MOCAT: A Metagenomics Assembly and Gene Prediction Toolkit. PLoS ONE. 2012;7(10):e47656. doi:10.1371/journal.pone.0047656. PMID:23082188. PMCID:PMC3474746.

Documentation

Downloads