dbCAN2

dbCAN2 annotates carbohydrate-active enzymes (CAZymes) in genomic and metagenomic datasets to characterize CAZyme repertoires involved in complex carbohydrate metabolism.


Key Features:

  • HMMER Search: Utilizes the dbCAN hidden Markov model (HMM) database to identify CAZyme sequences based on protein family HMMs.
  • DIAMOND Search: Performs rapid alignments against the pre-annotated CAZy sequence database to identify known CAZymes.
  • Hotpep Search: Searches against a curated database of conserved short peptides specific to CAZymes to enhance detection sensitivity.
  • Multi-tool consensus filtering: Combines outputs from HMMER, DIAMOND, and Hotpep and filters out CAZymes detected by only one method to improve annotation accuracy.
  • Nucleotide sequence input: Accepts nucleotide sequences directly for CAZyme annotation.
  • CAZyme Gene Cluster (CGC) prediction: Predicts physically linked CAZyme gene clusters to aid identification of putative polysaccharide utilization loci (PULs).

Scientific Applications:

  • Genome and metagenome CAZyme annotation: Annotating CAZymes within genomes and metagenomes to define CAZomes.
  • Plant and microbial community mining: Mining plant genomes and associated microbial communities for CAZymes involved in plant carbohydrate metabolism.
  • CAZyme gene cluster and PUL discovery: Identifying CAZyme gene clusters and putative polysaccharide utilization loci in microbial genomes or metagenomes.
  • Biomass and feedstock studies: Characterizing enzymes relevant to complex carbohydrate degradation for research on renewable energy, materials, and biological utilization of plant biomass.

Methodology:

Computational methods explicitly include HMMER searches using the dbCAN HMM database, DIAMOND searches against the pre-annotated CAZy sequence database, Hotpep searches against a curated conserved short-peptide database, combination of outputs with filtering of single-tool hits, and prediction of physically linked CAZyme gene clusters (CGCs).

Topics

Collections

Details

Tool Type:
web application
Added:
7/2/2018
Last Updated:
3/13/2019

Operations

Data Inputs & Outputs

Other operations do not define inputs or outputs.

Publications

Zhang H, Yohe T, Huang L, Entwistle S, Wu P, Yang Z, Busk PK, Xu Y, Yin Y. dbCAN2: a meta server for automated carbohydrate-active enzyme annotation. Nucleic Acids Research. 2018;46(W1):W95-W101. doi:10.1093/nar/gky418. PMID:29771380. PMCID:PMC6031026.

Funding: - National Science Foundation: DBI-1652164 - National Institutes of Health: 1R15GM114706 - Research & Artistry Award of the NIU: 2017-YIN - National Natural Science Foundation of China: 31728013

Documentation