Tallymer
Tallymer performs k-mer counting and indexing using enhanced suffix arrays to enable efficient analysis of repetitive genomic sequences, including transposable elements (TEs).
Key Features:
- Memory efficiency: Implements algorithms optimized for low memory usage to handle large genomic datasets.
- Flexibility in k-mer size: Supports a broad range of k-mer sizes for varied analytical resolutions.
- Large-data processing: Capable of processing datasets comprising several billion bases.
- Enhanced suffix arrays: Utilizes enhanced suffix arrays as the core data structure for indexing and counting.
- K-mer occurrence counting: Counts k-mer occurrences to distinguish repetitive transposable elements from low-copy genic regions.
- Low-coverage analysis: Enables effective k-mer frequency analysis at low sequencing coverage (e.g., ~0.45x).
Scientific Applications:
- Repeat annotation: Uses k-mer counts to annotate repeats and discriminate TEs from low-copy genomic regions.
- Genome evolution studies: Analyzes TE frequency and distribution to inform studies of genome evolution.
- Transposon activity analysis: Detects active transposons within expressed sequence tags (ESTs), as demonstrated in maize.
- Cross-cultivar comparisons: Identifies conserved retrotransposons across maize cultivars such as Mo17 and McC despite genomic rearrangements.
- Genome annotation support: Detects transposon-encoded genes with sensitivity and specificity in bacterial artificial chromosomes (BACs).
- Applied species: Has been applied to plant genomes including maize (B73), sorghum, and rice.
Methodology:
Uses enhanced suffix arrays for k-mer counting and indexing to enable efficient processing of large genomic datasets and k-mer frequency analysis at low coverage (e.g., ~0.45x).
Topics
Details
- Tool Type:
- workflow
- Operating Systems:
- Linux
- Programming Languages:
- Perl
- Added:
- 1/13/2017
- Last Updated:
- 11/25/2024
Operations
Publications
Kurtz S, Narechania A, Stein JC, Ware D. A new method to compute K-mer frequencies and its application to annotate large repetitive plant genomes. BMC Genomics. 2008;9(1). doi:10.1186/1471-2164-9-517. PMID:18976482. PMCID:PMC2613927.