findkm

findkm identifies and extracts k-mers (subsequences of length k) from genomic and proteomic sequences to quantify their frequency and distribution for sequence analysis.


Key Features:

  • Integration with EMBOSS: Integrates with the EMBOSS suite to interoperate with other EMBOSS applications.
  • EMBOSS C libraries: Builds on EMBOSS's C programming libraries to provide extensibility and integration.
  • K-mer extraction and analysis: Extracts all possible k-mers from input sequences and computes their frequency and distribution.
  • Sequence types supported: Operates on genomic (nucleotide) and proteomic (protein) sequence datasets.
  • Customization via ACD: Supports customization through ACD (Application Configuration Description) files.

Scientific Applications:

  • Genomic Research: Detecting motifs or sequence patterns in genomic data using k-mer analysis.
  • Proteomics: Identifying conserved regions or functional domains in protein sequences via k-mer distributions.
  • Comparative Genomics: Comparing k-mer distributions across species or strains to study evolutionary relationships.

Methodology:

findkm scans input sequences to extract all subsequences of a specified length k and analyzes k-mer frequency and distribution.

Topics

Collections

Details

License:
GPL-3.0
Maturity:
Mature
Cost:
Free of charge
Tool Type:
command-line tool
Operating Systems:
Linux, Windows, Mac
Programming Languages:
C
Added:
11/8/2015
Last Updated:
12/10/2018

Operations

Publications

Bleasby AJ, Ison JC, Rice PM. EMBOSS Administrator's Guide. Unknown Journal. 2011. doi:10.1017/cbo9781139151399.

Ison JC, Rice PM, Bleasby AJ. EMBOSS Developer's Guide. Unknown Journal. 2011. doi:10.1017/cbo9781139151405.

Rice P, Longden I, Bleasby A. EMBOSS: The European Molecular Biology Open Software Suite. Trends in Genetics. 2000;16(6):276-277. doi:10.1016/s0168-9525(00)02024-2.

Documentation

Downloads

Links