MOCAT
MOCAT processes shotgun metagenomics Illumina sequencing data to perform quality control, assembly, gene prediction, and taxonomic and functional profiling of microbial communities.
Key Features:
- Modular architecture: Modular architecture permitting customization and interchange of processing components.
- Raw data handling: Read preprocessing of raw fastQ files including quality control to filter low-quality reads and contaminants.
- Assembly: Assembly of metagenomic reads into contigs using assembly programs optimized for high-throughput sequencing.
- Gene prediction: Prediction of protein-coding genes on assembled contigs using advanced gene prediction algorithms.
- Taxonomic and functional profiling: Taxonomic and functional profiling by mapping reads against reference databases, enabling identification of phylogenetic marker genes, extraction or removal of specific sequences, and abundance calculations.
- MOCAT2 reference catalogs: Automated generation and annotation of non-redundant reference catalogs built from pre-computed assignments across 18 databases covering diverse functional categories (MOCAT2).
- HPC integration: Integration with LSF, PBS, and SGE queuing systems for high-performance computing on UNIX-based platforms.
- Implementation: Implemented in Perl 5 and Python 2.7.
- Output formats: Produces multi-sheet Excel documents and queryable SQL databases summarizing statistics from each processing step.
- License: Distributed under the GPL3 open-source license.
- Benchmarking: Benchmarked on artificial, real, and simulated metagenomes with reported improvements in selected quality metrics.
Scientific Applications:
- Shotgun metagenomics: Processing and analysis of shotgun metagenomics datasets generated with Illumina sequencing to profile microbial community composition and function.
- Gene catalog construction: Assembly of reads and prediction of protein-coding genes to construct gene catalogs and reference databases.
- Functional annotation: Functional annotation and characterization of metagenomes using non-redundant catalogs annotated across 18 functional databases.
- Taxonomic profiling: Generation of taxonomic profiles via identification of phylogenetic marker genes and read mapping.
- Abundance estimation and sequence filtering: Abundance estimation of taxa and genes and removal or extraction of specific sequences.
- Method benchmarking: Benchmarking and validation of metagenomic workflows using artificial, real, and simulated metagenomes.
Methodology:
Computational steps explicitly include raw fastQ quality control and contaminant filtering, assembly of reads into contigs with assembly programs, protein-coding gene prediction on contigs, read mapping to reference databases for taxonomic and functional profiling (including abundance calculations and sequence extraction/removal), automated creation and annotation of non-redundant reference catalogs from pre-computed assignments across 18 databases (MOCAT2), and execution via LSF, PBS, or SGE on UNIX; implemented in Perl 5 and Python 2.7 with outputs as multi-sheet Excel and SQL databases.
Topics
Collections
Details
- License:
- GPL-3.0
- Maturity:
- Mature
- Cost:
- Free of charge
- Tool Type:
- command-line tool
- Operating Systems:
- Linux, Mac
- Programming Languages:
- Perl, Python
- Added:
- 4/5/2017
- Last Updated:
- 11/24/2024
Operations
Data Inputs & Outputs
Community profiling
Outputs
Publications
Kultima JR, Coelho LP, Forslund K, Huerta-Cepas J, Li SS, Driessen M, Voigt AY, Zeller G, Sunagawa S, Bork P. MOCAT2: a metagenomic assembly, annotation and profiling framework. Bioinformatics. 2016;32(16):2520-2523. doi:10.1093/bioinformatics/btw183. PMID:27153620. PMCID:PMC4978931.
Kultima JR, Sunagawa S, Li J, Chen W, Chen H, Mende DR, Arumugam M, Pan Q, Liu B, Qin J, Wang J, Bork P. MOCAT: A Metagenomics Assembly and Gene Prediction Toolkit. PLoS ONE. 2012;7(10):e47656. doi:10.1371/journal.pone.0047656. PMID:23082188. PMCID:PMC3474746.