SAKEIMA
SAKEIMA approximates frequencies of the most frequent k-mers in high-throughput sequencing datasets to enable efficient k-mer-based analyses.
Key Features:
- Sampling-Based Approach: Employs a sampling strategy to approximate frequent k-mers and their frequencies without processing entire datasets.
- Rigorous Quality Guarantees: Leverages statistical learning theory and the Vapnik-Chervonenkis (VC) dimension to provide theoretical bounds on approximation quality and sample size.
- Efficiency for Large Datasets: Processes only a fraction of the data to achieve substantial speedups in analyzing large-scale sequencing data.
- Accurate Distance Estimation: Produces k-mer frequency estimates that enable precise calculation of k-mer-based distances between datasets.
Scientific Applications:
- Metagenomics: Facilitates computation of k-mer-based distances among metagenomic datasets for comparative analyses of microbial communities.
- Genome Assembly and Annotation: Identifies frequent k-mers that can inform genome assembly steps and support functional annotation tasks.
- Comparative Genomics: Enables comparison of genomic sequences across species or strains by providing approximate k-mer frequency profiles.
Methodology:
Uses a sampling-based algorithm informed by statistical learning theory and the Vapnik-Chervonenkis (VC) dimension to determine practical sample size bounds and approximate frequent k-mer frequencies.
Topics
Details
- Programming Languages:
- C++
- Added:
- 1/18/2021
- Last Updated:
- 2/10/2021
Operations
Publications
Pellegrina L, Pizzi C, Vandin F. Fast Approximation of Frequent<i>k</i>-Mers and Applications to Metagenomics. Journal of Computational Biology. 2020;27(4):534-549. doi:10.1089/cmb.2019.0314. PMID:31891535.
PMID: 31891535