SANDPUMA

SANDPUMA predicts substrate specificities of nonribosomal peptide synthetase adenylation (A) domains from DNA sequences to support discovery and characterization of nonribosomally synthesized peptides (NRPs).


Key Features:

  • Ensemble algorithm: Integrates multiple high-performing algorithms into an ensemble to improve accuracy of A-domain substrate predictions.
  • prediCAT phylogenetics-inspired estimation: Incorporates prediCAT to quantitatively estimate the predictability of individual A-domains using phylogenetics-inspired principles.
  • Benchmarking on independent test set: Evaluated against an independent test set of 434 A-domain sequences, demonstrating that active-site-motif-focused algorithms outperform whole-domain methods.
  • antiSMASH integration: Integrates with the antiSMASH biosynthetic gene cluster analysis pipeline to combine genomic cluster annotation with A-domain substrate prediction.

Scientific Applications:

  • NRP chemical diversity exploration: Applied to 7635 Actinobacteria genomes to reveal greater nonribosomal peptide (NRP) diversity than previously estimated.
  • Natural product discovery and biotechnology: Guides identification of candidate A-domains and predicted substrates to inform natural product discovery and biotechnological exploitation.

Methodology:

High-performing algorithms were retrained on a comprehensive dataset and integrated into an ensemble; prediCAT provides per-domain predictability estimates; models were benchmarked against an independent test set of 434 A-domain sequences with comparisons between active-site-motif-focused and whole-domain methods.

Topics

Details

License:
GPL-3.0
Tool Type:
command-line tool
Operating Systems:
Linux, Mac
Programming Languages:
Perl, Python
Added:
6/12/2018
Last Updated:
11/25/2024

Operations

Publications

Chevrette MG, Aicheler F, Kohlbacher O, Currie CR, Medema MH. SANDPUMA: ensemble predictions of nonribosomal peptide chemistry reveal biosynthetic diversity across <i>Actinobacteria</i>. Bioinformatics. 2017;33(20):3202-3210. doi:10.1093/bioinformatics/btx400. PMID:28633438. PMCID:PMC5860034.

PMID: 28633438
PMCID: PMC5860034
Funding: - National Institutes of Health: U19 Al109673

Documentation