BiG-SCAPE
BiG-SCAPE constructs sequence similarity networks of biosynthetic gene clusters (BGCs) to cluster them into gene cluster families (GCFs) and quantify biosynthetic diversity across large genomic datasets.
Key Features:
- Sequence similarity networks: Constructs sequence similarity networks of BGCs and groups BGCs into gene cluster families (GCFs).
- Distance matrix computation: Rapidly computes a distance matrix between gene clusters by comparing protein domain content, domain order, domain copy number, and protein sequence identity.
- Composite distance metric: Implements a composite distance metric integrating the Jaccard Index for domain types, Domain Sequence Similarity, and Adjacency.
- Protein-domain comparisons: Performs detailed comparisons of protein domain content, order, copy number, and sequence identity among clusters.
- Integration with CORASON: Can be used in conjunction with CORASON (Core Analysis of Syntenic Orthologues to Prioritize Natural Product Gene Clusters) to relate BGC networks to enzyme phylogenies.
- Metabolomic validation: Validated by correlation with metabolomic data from 363 actinobacterial strains.
- Natural product discovery: Enabled mapping of detoxin/rimosamide-related GCFs and the characterization of seven novel detoxin analogues.
- Implementation: Implemented in Python.
Scientific Applications:
- Genome-scale BGC mining: Mining biosynthetic gene clusters across large numbers of genomes to assess natural product diversity.
- GCF-based diversity analysis: Grouping BGCs into GCFs to study the distribution and diversity of biosynthetic pathways.
- Genomics–metabolomics integration: Correlating BGC sequence similarity networks with metabolomic data (e.g., 363 actinobacterial strains) to associate gene clusters with metabolites.
- Phylogenetic prioritization: Prioritizing natural product gene clusters and exploring enzyme phylogenies when combined with CORASON.
- Natural product discovery: Facilitating discovery and characterization of novel natural products such as seven detoxin analogues from detoxin/rimosamide-related GCFs.
Methodology:
Calculates a distance matrix by comparing protein domain content, domain order, domain copy number, and protein sequence identity; computes a composite distance metric integrating the Jaccard Index for domain types, Domain Sequence Similarity, and Adjacency; constructs sequence similarity networks and clusters BGCs into GCFs; implemented in Python; can be used with CORASON to analyze syntenic orthologues and enzyme phylogenies.
Topics
Details
- Tool Type:
- command-line tool
- Programming Languages:
- Python
- Added:
- 1/14/2020
- Last Updated:
- 11/24/2024
Operations
Data Inputs & Outputs
Clustering
Publications
Navarro-Muñoz JC, Selem-Mojica N, Mullowney MW, Kautsar SA, Tryon JH, Parkinson EI, De Los Santos ELC, Yeong M, Cruz-Morales P, Abubucker S, Roeters A, Lokhorst W, Fernandez-Guerra A, Cappelini LTD, Goering AW, Thomson RJ, Metcalf WW, Kelleher NL, Barona-Gomez F, Medema MH. A computational framework to explore large-scale biosynthetic diversity. Nature Chemical Biology. 2019;16(1):60-68. doi:10.1038/s41589-019-0400-9. PMID:31768033. PMCID:PMC6917865.