K2Mem

K2Mem enhances metagenomic read classification by discovering and incorporating discriminative k-mers from input sequencing data to improve taxonomic annotation of sequencing reads.


Key Features:

  • Discriminative k-mer discovery: Identifies novel k-mers directly from input sequencing reads that distinguish between different taxa.
  • Reference k-mer library augmentation: Augments existing reference k-mer libraries with discovered discriminative k-mers.
  • Improved classification metrics: Increases recall and F-measure while maintaining high precision in metagenomic classification.
  • Robustness to reference divergence: Addresses sequence variation and highly mutated genomes by incorporating sample-derived k-mers when closely related reference genomes are unavailable.
  • Empirical evaluation: Performance has been evaluated under multiple conditions and compared against existing metagenomic classification tools.

Scientific Applications:

  • Taxonomic annotation of metagenomic reads: Improves assignment of sequencing reads to taxa in complex samples.
  • Species identification in environmental and clinical samples: Supports detection of species where reference genomes may be incomplete or divergent.
  • Analysis of rapidly evolving organisms: Enhances classification of highly mutated viral genomes and other rapidly evolving taxa.
  • Metagenomic sensitivity and specificity improvement: Increases sensitivity without compromising specificity in metagenomic studies.

Methodology:

Analyzes input sequencing reads to discover unique discriminative k-mers and augments the existing reference k-mer library with those k-mers.

Topics

Details

License:
GPL-3.0
Cost:
Free of charge
Tool Type:
command-line tool
Operating Systems:
Mac, Linux, Windows
Programming Languages:
C++, Perl, Shell
Added:
4/26/2022
Last Updated:
4/26/2022

Operations

Publications

Storato D, Comin M. K2Mem: Discovering Discriminative K-mers From Sequencing Data for Metagenomic Reads Classification. IEEE/ACM Transactions on Computational Biology and Bioinformatics. 2022;19(1):220-229. doi:10.1109/tcbb.2021.3117406. PMID:34606462.