hash-cgMLST

hash-cgMLST converts core genome multilocus sequence typing (cgMLST) alleles into unique hashes to enable rapid, database-free comparison of bacterial genomes for genotyping and transmission analysis of pathogens such as Clostridium difficile.


Key Features:

  • Hash-based allele representation: Alleles at each cgMLST gene are converted into unique hash strings, removing the need for a centralized, sequentially numbered allele database.
  • Reproducibility and efficiency: Maintains cgMLST reproducibility while allowing rapid processing and comparison of hash profiles with minimal performance penalty.
  • Discriminatory power: Demonstrates robust discrimination of closely related genomes and accurate identification of replicate sequence pairs with minimal gene differences in Clostridium difficile datasets.
  • High-throughput comparison: Enables comparison of a single genome against up to 100,000 others in under one minute.
  • Resource profile: The hash-based comparison step requires minimal system resources, although genome assembly using SPAdes is resource-intensive.
  • Computational environment compatibility: Has been run on MacOS, Ubuntu, and CentOS and is compatible with Java 8+ and Nextflow workflows.

Scientific Applications:

  • Pathogen whole-genome sequencing: Rapidly identifies related genomes to support investigation of infection transmission dynamics, exemplified by Clostridium difficile studies.
  • Epidemiology and public health surveillance: Enables detection of clusters of related isolates for outbreak investigation and large-scale surveillance.
  • Preliminary large-scale screening: Serves as a fast initial genotyping approach, with mapping-based methods recommended for follow-up to define precise genetic relationships due to a higher potential for false variants.

Methodology:

Alleles are converted into unique hash strings derived from genome assemblies (assemblies generated using SPAdes), hash profiles are compared for cgMLST typing, and performance was validated using repeated sequencing of Clostridium difficile isolates and datasets of consecutive infection isolates from multiple hospitals.

Topics

Details

Tool Type:
command-line tool
Programming Languages:
R, Python
Added:
1/14/2020
Last Updated:
12/7/2020

Operations

Publications

Eyre DW, Peto TEA, Crook DW, Walker AS, Wilcox MH. Hash-Based Core Genome Multilocus Sequence Typing for Clostridium difficile. Journal of Clinical Microbiology. 2019;58(1). doi:10.1128/jcm.01037-19. PMID:31666367. PMCID:PMC6935933.

PMID: 31666367
PMCID: PMC6935933
Funding: - National Institute for Health Research: HPRU in Healthcare Associated Infections and AMR