SAKEIMA

SAKEIMA approximates frequencies of the most frequent k-mers in high-throughput sequencing datasets to enable efficient k-mer-based analyses.


Key Features:

  • Sampling-Based Approach: Employs a sampling strategy to approximate frequent k-mers and their frequencies without processing entire datasets.
  • Rigorous Quality Guarantees: Leverages statistical learning theory and the Vapnik-Chervonenkis (VC) dimension to provide theoretical bounds on approximation quality and sample size.
  • Efficiency for Large Datasets: Processes only a fraction of the data to achieve substantial speedups in analyzing large-scale sequencing data.
  • Accurate Distance Estimation: Produces k-mer frequency estimates that enable precise calculation of k-mer-based distances between datasets.

Scientific Applications:

  • Metagenomics: Facilitates computation of k-mer-based distances among metagenomic datasets for comparative analyses of microbial communities.
  • Genome Assembly and Annotation: Identifies frequent k-mers that can inform genome assembly steps and support functional annotation tasks.
  • Comparative Genomics: Enables comparison of genomic sequences across species or strains by providing approximate k-mer frequency profiles.

Methodology:

Uses a sampling-based algorithm informed by statistical learning theory and the Vapnik-Chervonenkis (VC) dimension to determine practical sample size bounds and approximate frequent k-mer frequencies.

Topics

Details

Programming Languages:
C++
Added:
1/18/2021
Last Updated:
2/10/2021

Operations

Publications

Pellegrina L, Pizzi C, Vandin F. Fast Approximation of Frequent<i>k</i>-Mers and Applications to Metagenomics. Journal of Computational Biology. 2020;27(4):534-549. doi:10.1089/cmb.2019.0314. PMID:31891535.