CDKAM

CDKAM classifies metagenomic sequences from third-generation sequencing (TGS) data using discriminative k-mers and an approximate matching algorithm to achieve species-level taxonomic identification despite high TGS error rates.


Key Features:

  • Discriminative k-mers: Uses discriminative k-mers to represent taxon-specific sequence signatures for metagenomic classification.
  • Approximate matching strategy: Employs an approximate matching algorithm to accommodate the high error rates and error patterns characteristic of third-generation sequencing (TGS) when identifying k-mers.
  • Two-phase approach: Implements a rapid quick mapping phase to identify candidate k-mer matches followed by a dynamic programming refinement phase to improve match accuracy.
  • Performance characteristics: Evaluated on simulated and real TGS datasets, CDKAM outperforms existing methods for sequences ~1000–1500 bases, achieving higher species-level accuracy with reduced memory use.

Scientific Applications:

  • Metagenomic taxonomic classification of TGS data: Classifies metagenome sequences generated by third-generation sequencing technologies to enable taxonomic profiling.
  • Species-level identification in microbial community studies: Provides high species-level accuracy useful for analyses of complex microbial communities.

Methodology:

CDKAM applies discriminative k-mers with an approximate matching algorithm implemented as a two-phase process consisting of a quick mapping phase followed by dynamic programming refinement.

Topics

Details

Programming Languages:
C++, Perl, Shell
Added:
1/18/2021
Last Updated:
2/10/2021

Operations

Publications

Bui V, Wei C. CDKAM: a taxonomic classification tool using discriminative k-mers and approximate matching strategies. BMC Bioinformatics. 2020;21(1). doi:10.1186/s12859-020-03777-y. PMID:33081690. PMCID:PMC7576720.

PMID: 33081690
PMCID: PMC7576720
Funding: - Cross-Institute Research Fund of Shanghai Jiao Tong University: YG2017ZD01 - National Natural Science Foundation of China: 61472246 - National Basic Research Program of China: 2013CB956103