SANDPUMA
SANDPUMA predicts substrate specificities of nonribosomal peptide synthetase adenylation (A) domains from DNA sequences to support discovery and characterization of nonribosomally synthesized peptides (NRPs).
Key Features:
- Ensemble algorithm: Integrates multiple high-performing algorithms into an ensemble to improve accuracy of A-domain substrate predictions.
- prediCAT phylogenetics-inspired estimation: Incorporates prediCAT to quantitatively estimate the predictability of individual A-domains using phylogenetics-inspired principles.
- Benchmarking on independent test set: Evaluated against an independent test set of 434 A-domain sequences, demonstrating that active-site-motif-focused algorithms outperform whole-domain methods.
- antiSMASH integration: Integrates with the antiSMASH biosynthetic gene cluster analysis pipeline to combine genomic cluster annotation with A-domain substrate prediction.
Scientific Applications:
- NRP chemical diversity exploration: Applied to 7635 Actinobacteria genomes to reveal greater nonribosomal peptide (NRP) diversity than previously estimated.
- Natural product discovery and biotechnology: Guides identification of candidate A-domains and predicted substrates to inform natural product discovery and biotechnological exploitation.
Methodology:
High-performing algorithms were retrained on a comprehensive dataset and integrated into an ensemble; prediCAT provides per-domain predictability estimates; models were benchmarked against an independent test set of 434 A-domain sequences with comparisons between active-site-motif-focused and whole-domain methods.
Topics
Details
- License:
- GPL-3.0
- Tool Type:
- command-line tool
- Operating Systems:
- Linux, Mac
- Programming Languages:
- Perl, Python
- Added:
- 6/12/2018
- Last Updated:
- 11/25/2024
Operations
Publications
Chevrette MG, Aicheler F, Kohlbacher O, Currie CR, Medema MH. SANDPUMA: ensemble predictions of nonribosomal peptide chemistry reveal biosynthetic diversity across <i>Actinobacteria</i>. Bioinformatics. 2017;33(20):3202-3210. doi:10.1093/bioinformatics/btx400. PMID:28633438. PMCID:PMC5860034.