findkm
findkm identifies and extracts k-mers (subsequences of length k) from genomic and proteomic sequences to quantify their frequency and distribution for sequence analysis.
Key Features:
- Integration with EMBOSS: Integrates with the EMBOSS suite to interoperate with other EMBOSS applications.
- EMBOSS C libraries: Builds on EMBOSS's C programming libraries to provide extensibility and integration.
- K-mer extraction and analysis: Extracts all possible k-mers from input sequences and computes their frequency and distribution.
- Sequence types supported: Operates on genomic (nucleotide) and proteomic (protein) sequence datasets.
- Customization via ACD: Supports customization through ACD (Application Configuration Description) files.
Scientific Applications:
- Genomic Research: Detecting motifs or sequence patterns in genomic data using k-mer analysis.
- Proteomics: Identifying conserved regions or functional domains in protein sequences via k-mer distributions.
- Comparative Genomics: Comparing k-mer distributions across species or strains to study evolutionary relationships.
Methodology:
findkm scans input sequences to extract all subsequences of a specified length k and analyzes k-mer frequency and distribution.
Topics
Collections
Details
- License:
- GPL-3.0
- Maturity:
- Mature
- Cost:
- Free of charge
- Tool Type:
- command-line tool
- Operating Systems:
- Linux, Windows, Mac
- Programming Languages:
- C
- Added:
- 11/8/2015
- Last Updated:
- 12/10/2018
Operations
Publications
Bleasby AJ, Ison JC, Rice PM. EMBOSS Administrator's Guide. Unknown Journal. 2011. doi:10.1017/cbo9781139151399.
Ison JC, Rice PM, Bleasby AJ. EMBOSS Developer's Guide. Unknown Journal. 2011. doi:10.1017/cbo9781139151405.
Rice P, Longden I, Bleasby A. EMBOSS: The European Molecular Biology Open Software Suite. Trends in Genetics. 2000;16(6):276-277. doi:10.1016/s0168-9525(00)02024-2.
Documentation
Terms of use
http://emboss.open-bio.org/html/dev/ch01s01.htmlCitation instructions
http://emboss.open-bio.org/html/use/pr02s04.html