BAGEL2

BAGEL2 identifies putative bacteriocin genes and annotates their genomic contexts in bacterial genomes, unfinished genome assemblies, and metagenomic datasets to enable discovery and characterization of antimicrobial peptides.


Key Features:

  • Conserved domains and physical properties: Identifies potential bacteriocins using conserved sequence domains and characteristic physical properties of antimicrobial peptides.
  • Genomic context analysis: Detects biosynthesis, transport, and immunity genes to evaluate the genomic neighborhoods of putative bacteriocin genes.
  • High-throughput processing: Supports large-scale mining of complete and unfinished genomes as well as metagenomic samples.
  • Parameter-free, class-specific mining: Performs class-specific bacteriocin mining without requiring user-defined parameters.
  • Novel hidden Markov models (HMMs): Applies novel HMMs for prediction of bacteriocin sub-classes, interpreted using simple decision rules.
  • Automated genetic context annotation: Annotates genetic contexts automatically using combinations of PFAM domains and databases of known context genes.
  • Expert-validated database: Integrates an expert-validated database of bacteriocins for reference and comparison.
  • Fine-tuned scoring system: Uses a scoring system refined with expert knowledge derived from screening all bacterial genomes available at NCBI.

Scientific Applications:

  • Genome-scale bacteriocin discovery: Enables comprehensive mining of bacteriocins across diverse bacterial species and genome assemblies.
  • Metagenomic discovery of novel antimicrobials: Facilitates identification of novel bacteriocins within metagenomic datasets.
  • Functional and regulatory inference: Uses genomic context (biosynthesis, transport, immunity genes) to infer regulatory and functional aspects of bacteriocin production.
  • Large-scale screening and curation: Supports large-scale surveys and curation workflows by combining predictive models with an expert-validated bacteriocin database and tuned scoring.

Methodology:

Identification based on conserved domains and physical properties; prediction of subclasses using novel HMMs interpreted by simple decision rules; automated annotation of genetic context using PFAM domain combinations and databases of known context genes; scoring refined by screening all bacterial genomes in NCBI.

Topics

Details

Tool Type:
web application
Added:
3/25/2017
Last Updated:
11/25/2024

Operations

Data Inputs & Outputs

Publications

de Jong A, van Heel AJ, Kok J, Kuipers OP. BAGEL2: mining for bacteriocins in genomic data. Nucleic Acids Research. 2010;38(suppl_2):W647-W651. doi:10.1093/nar/gkq365. PMID:20462861. PMCID:PMC2896169.