KmerGO
KmerGO identifies group-specific k-mers and trait-associated sequences in high-throughput sequencing datasets to characterize genetic and microbial differences between sample groups.
Key Features:
- Identification of Group-Specific k-mers: Detects group-specific k-mers up to 40 base pairs (bps) to capture single nucleotide variants (SNVs), gene families, microbial species, and other group-unique elements.
- Efficiency and Resource Management: Processes large .fasta datasets (example: 1.05 TB) with low memory footprint (≤1 GB) and an example runtime of ~21.5 hours on a standard workstation.
- Trait-Associated Sequence Capture: Identifies sequences associated with continuous-trait datasets for discovery of trait-linked markers.
- Parallel Computing: Implements multi-process parallel computing to accelerate k-mer discovery and group comparisons.
- Integration with Downstream Tools: Produces group-specific k-mers and sequences suitable for downstream analyses such as biomarker discovery, genetic variant analysis, species identification, and gene characterization.
Scientific Applications:
- Comparative Genomics: Enables identification of group-specific markers for comparative genomics studies.
- Microbial Ecology: Supports detection of microbial species and signatures that distinguish ecological groups in metagenomic datasets.
- Personalized Medicine: Facilitates discovery of genetic markers and variants relevant to personalized medicine.
- Biomarker Discovery: Provides candidate k-mer biomarkers for disease association and diagnostic development.
- Evolutionary Biology: Assists in detecting sequence differences relevant to evolutionary analyses.
- Environmental Microbiology: Supports identification of environment-specific microbial signatures from metagenomic data.
- Trait Association Studies: Enables capture of sequences associated with continuous traits for studies linking sequence variation to trait variation.
Methodology:
Performs k-mer–based identification (k ≤ 40 bp) by comparing presence and abundance between groups on .fasta input, captures SNVs, gene families, microbial species and sequences associated with continuous-trait datasets, uses multi-process parallel computing, and outputs group-specific k-mers/sequences for downstream analysis.
Topics
Details
- Tool Type:
- command-line tool, desktop application
- Programming Languages:
- Python
- Added:
- 1/18/2021
- Last Updated:
- 2/12/2021
Operations
Publications
Wang Y, Chen Q, Deng C, Zheng Y, Sun F. KmerGO: A Tool to Identify Group-Specific Sequences With k-mers. Frontiers in Microbiology. 2020;11. doi:10.3389/fmicb.2020.02067. PMID:32983048. PMCID:PMC7477287.