BiG-SCAPE

BiG-SCAPE constructs sequence similarity networks of biosynthetic gene clusters (BGCs) to cluster them into gene cluster families (GCFs) and quantify biosynthetic diversity across large genomic datasets.


Key Features:

  • Sequence similarity networks: Constructs sequence similarity networks of BGCs and groups BGCs into gene cluster families (GCFs).
  • Distance matrix computation: Rapidly computes a distance matrix between gene clusters by comparing protein domain content, domain order, domain copy number, and protein sequence identity.
  • Composite distance metric: Implements a composite distance metric integrating the Jaccard Index for domain types, Domain Sequence Similarity, and Adjacency.
  • Protein-domain comparisons: Performs detailed comparisons of protein domain content, order, copy number, and sequence identity among clusters.
  • Integration with CORASON: Can be used in conjunction with CORASON (Core Analysis of Syntenic Orthologues to Prioritize Natural Product Gene Clusters) to relate BGC networks to enzyme phylogenies.
  • Metabolomic validation: Validated by correlation with metabolomic data from 363 actinobacterial strains.
  • Natural product discovery: Enabled mapping of detoxin/rimosamide-related GCFs and the characterization of seven novel detoxin analogues.
  • Implementation: Implemented in Python.

Scientific Applications:

  • Genome-scale BGC mining: Mining biosynthetic gene clusters across large numbers of genomes to assess natural product diversity.
  • GCF-based diversity analysis: Grouping BGCs into GCFs to study the distribution and diversity of biosynthetic pathways.
  • Genomics–metabolomics integration: Correlating BGC sequence similarity networks with metabolomic data (e.g., 363 actinobacterial strains) to associate gene clusters with metabolites.
  • Phylogenetic prioritization: Prioritizing natural product gene clusters and exploring enzyme phylogenies when combined with CORASON.
  • Natural product discovery: Facilitating discovery and characterization of novel natural products such as seven detoxin analogues from detoxin/rimosamide-related GCFs.

Methodology:

Calculates a distance matrix by comparing protein domain content, domain order, domain copy number, and protein sequence identity; computes a composite distance metric integrating the Jaccard Index for domain types, Domain Sequence Similarity, and Adjacency; constructs sequence similarity networks and clusters BGCs into GCFs; implemented in Python; can be used with CORASON to analyze syntenic orthologues and enzyme phylogenies.

Topics

Details

Tool Type:
command-line tool
Programming Languages:
Python
Added:
1/14/2020
Last Updated:
11/24/2024

Operations

Data Inputs & Outputs

Publications

Navarro-Muñoz JC, Selem-Mojica N, Mullowney MW, Kautsar SA, Tryon JH, Parkinson EI, De Los Santos ELC, Yeong M, Cruz-Morales P, Abubucker S, Roeters A, Lokhorst W, Fernandez-Guerra A, Cappelini LTD, Goering AW, Thomson RJ, Metcalf WW, Kelleher NL, Barona-Gomez F, Medema MH. A computational framework to explore large-scale biosynthetic diversity. Nature Chemical Biology. 2019;16(1):60-68. doi:10.1038/s41589-019-0400-9. PMID:31768033. PMCID:PMC6917865.

PMID: 31768033
PMCID: PMC6917865
Funding: - Consejo Nacional de Ciencia y Tecnología: 2017_051TAMU, CBS2017_285746, PhD scholarship 204482, postdoctoral scholarship 263661 - U.S. Department of Health & Human Services | NIH | National Institute of General Medical Sciences: F32GM120999 - U.S. Department of Health & Human Services | NIH | National Cancer Institute: F32CA221327 - Nederlandse Organisatie voor Wetenschappelijk Onderzoek: 863.15.002 - U.S. Department of Health & Human Services | National Institutes of Health: U01GM110706 - RCUK | Biotechnology and Biological Sciences Research Council: BB/M017982/1 - EC | Horizon 2020 Framework Programme: 634486 - U.S. Department of Health & Human Services | NIH | National Center for Complementary and Integrative Health: R01AT009143

Links