WhatsGNU

WhatsGNU identifies proteomic novelty by using exact-match proteomic compression to rapidly classify genomes, quantify allelic novelty with the GNU score, and contextualize individual genomes within large collections for genomic diversity analyses.


Key Features:

  • Exact-match proteomic compression: Compresses proteomes by identifying exact-match protein sequences across genomes to reduce redundancy.
  • GNU score: Quantifies proteomic novelty by ranking protein sequences based on observed exact matches across known genomes within a species.
  • Use of public-database variation: Leverages natural variation present in public databases to contextualize and rank protein sequences.
  • Rapid genome classification: Classifies new genomes swiftly by comparing their proteomes to the compressed reference of exact-match proteins.
  • Whole-protein and allele reports: Generates comprehensive whole-protein reports that list protein alleles and highlight potential novel functional differences.
  • Scalability: Employs computationally scalable methodologies to process large collections of genome sequences efficiently.

Scientific Applications:

  • Panallelome characterization: Characterizes total allelic diversity (panallelome) of bacterial species, including Salmonella enterica, Mycobacterium tuberculosis, Pseudomonas aeruginosa, and Staphylococcus aureus.
  • Microbial genomics: Supports analyses of genetic variation relevant to disease mechanisms, antibiotic resistance, and vaccine development.

Methodology:

Performs exact-match proteomic compression, ranks protein sequences by observed exact matches in public genomes to compute a GNU score, classifies genomes based on these comparisons, and outputs whole-protein allele reports.

Topics

Details

License:
GPL-3.0
Programming Languages:
Python
Added:
1/18/2021
Last Updated:
3/14/2021

Operations

Publications

Moustafa AM, Planet PJ. WhatsGNU: a tool for identifying proteomic novelty. Genome Biology. 2020;21(1). doi:10.1186/s13059-020-01965-w. PMID:32138767. PMCID:PMC7059281.

PMID: 32138767
PMCID: PMC7059281
Funding: - National Institutes of Health: 1K08AI101005, 1R01AI137526-01