KmerGO

KmerGO identifies group-specific k-mers and trait-associated sequences in high-throughput sequencing datasets to characterize genetic and microbial differences between sample groups.


Key Features:

  • Identification of Group-Specific k-mers: Detects group-specific k-mers up to 40 base pairs (bps) to capture single nucleotide variants (SNVs), gene families, microbial species, and other group-unique elements.
  • Efficiency and Resource Management: Processes large .fasta datasets (example: 1.05 TB) with low memory footprint (≤1 GB) and an example runtime of ~21.5 hours on a standard workstation.
  • Trait-Associated Sequence Capture: Identifies sequences associated with continuous-trait datasets for discovery of trait-linked markers.
  • Parallel Computing: Implements multi-process parallel computing to accelerate k-mer discovery and group comparisons.
  • Integration with Downstream Tools: Produces group-specific k-mers and sequences suitable for downstream analyses such as biomarker discovery, genetic variant analysis, species identification, and gene characterization.

Scientific Applications:

  • Comparative Genomics: Enables identification of group-specific markers for comparative genomics studies.
  • Microbial Ecology: Supports detection of microbial species and signatures that distinguish ecological groups in metagenomic datasets.
  • Personalized Medicine: Facilitates discovery of genetic markers and variants relevant to personalized medicine.
  • Biomarker Discovery: Provides candidate k-mer biomarkers for disease association and diagnostic development.
  • Evolutionary Biology: Assists in detecting sequence differences relevant to evolutionary analyses.
  • Environmental Microbiology: Supports identification of environment-specific microbial signatures from metagenomic data.
  • Trait Association Studies: Enables capture of sequences associated with continuous traits for studies linking sequence variation to trait variation.

Methodology:

Performs k-mer–based identification (k ≤ 40 bp) by comparing presence and abundance between groups on .fasta input, captures SNVs, gene families, microbial species and sequences associated with continuous-trait datasets, uses multi-process parallel computing, and outputs group-specific k-mers/sequences for downstream analysis.

Topics

Details

Tool Type:
command-line tool, desktop application
Programming Languages:
Python
Added:
1/18/2021
Last Updated:
2/12/2021

Operations

Publications

Wang Y, Chen Q, Deng C, Zheng Y, Sun F. KmerGO: A Tool to Identify Group-Specific Sequences With k-mers. Frontiers in Microbiology. 2020;11. doi:10.3389/fmicb.2020.02067. PMID:32983048. PMCID:PMC7477287.