KEC
KEC identifies unique amino acid and nucleic acid sequences specific to a target organism by eliminating k-mers found in non-target sequences within large-scale sequence datasets.
Key Features:
- K-mer-based filtering: Uses a k-mer-based elimination approach to remove sequences shared with non-target datasets.
- Cross-reference elimination: Cross-references input sequences against a comprehensive database of non-target sequences to isolate candidate unique sequences.
- Sequence type support: Operates on both amino acid and nucleic acid sequence data.
- Genome-scale performance: Performs rapid processing on datasets comparable in size to whole genomes.
Scientific Applications:
- Genomic Research: Identifies species-specific genetic markers to support studies in evolution, taxonomy, and biodiversity.
- Metagenomics: Differentiates genomic content of organisms within mixed samples to aid metagenomic analyses.
- Comparative Genomics: Enables identification of sequences that distinguish one organism from another for evolutionary and functional annotation studies.
Methodology:
KEC employs a k-mer-based approach to eliminate non-target sequences, focusing on the identification of unique sequences within the target dataset; it cross-references input sequences against a comprehensive database of non-target sequences to narrow down candidates exclusive to the organism of interest.
Topics
Details
- Tool Type:
- command-line tool
- Added:
- 10/4/2021
- Last Updated:
- 10/4/2021
Operations
Publications
Beran P, Stehlíková D, Cohen SP, Čurn V. KEC: unique sequence search by K-mer exclusion. Bioinformatics. 2021;37(19):3349-3350. doi:10.1093/bioinformatics/btab196. PMID:33755102.
PMID: 33755102
Funding: - Ministry of Education, Youths, and Sports: MSMT-15739/2019-8
- European Cooperation in Science and Technology: CA16107
- AFRI Education and Workforce Development Postdoctoral Fellowship: 2018-08122
Documentation
Links
Issue tracker
https://github.com/berybox/KEC/issues