Kmerator
Kmerator predicts and analyzes specific k-mers from RNA-sequencing (RNA-seq) data to enable k-mer based transcript quantification and signature extraction using a reference genome and an ENSEMBL-like transcriptome.
Key Features:
- Specific k-mer prediction: Predicts and analyzes specific k-mers (tags) derived from input sequences for gene- and transcript-level signatures.
- Contig generation: Generates contigs as sequences of consecutive k-mers overlapping by k-1 from specific k-mers.
- Jellyfish indexing: Uses Jellyfish to build two requestable indexes from the reference genome and the ENSEMBL-like transcriptome.
- K-mer decomposition and counting: Decomposes input transcript or gene sequences into k-mers and counts occurrences within both the genome and the transcriptome.
- Human gene coverage and accuracy: Produces specific k-mers for 97% of human genes and enables accurate gene expression quantification in simulated datasets.
- Suite composition: Forms part of the Kmerator Suite comprising three tools that extract, quantify, and analyze k-mer signatures, including KmerExploR.
- Metadata inference (KmerExploR): Applies gene-specific k-mer predictors to infer metadata such as library protocol, sample features, and potential contaminations.
- Advanced queries: Supports queries for known or novel biomarkers, including identification of mutations, gene fusions, and long non-coding RNAs.
- Computational efficiency: Provides a k-mer decomposition alternative to alignment-based methods to reduce computational resource requirements for large-scale RNA-seq analyses.
Scientific Applications:
- Gene expression quantification: Quantifies transcript- and gene-level expression from RNA-seq via specific k-mer counts.
- Sample metadata inference: Infers library protocol, sample features, and contaminations from RNA-seq datasets using KmerExploR.
- Biomarker discovery: Enables queries for known or novel biomarkers from k-mer signatures.
- Detection of genetic events: Identifies mutations and gene fusions from k-mer patterns.
- Long non-coding RNA analysis: Detects long non-coding RNAs using k-mer based signatures.
- Large-scale RNA-seq processing: Facilitates resource-efficient analysis of large RNA-seq cohorts using k-mer decomposition instead of alignment.
Methodology:
Uses Jellyfish to create two requestable indexes from the reference genome and the ENSEMBL-like transcriptome; decomposes input transcript or gene sequences into k-mers and counts their occurrences within genome and transcriptome; generates contigs by joining consecutive specific k-mers overlapping by k-1; predicts specific k-mers (tags) derived from input sequences.
Topics
Details
- License:
- GPL-3.0
- Maturity:
- Emerging
- Cost:
- Free of charge
- Tool Type:
- command-line tool
- Programming Languages:
- Python
- Added:
- 10/4/2021
- Last Updated:
- 11/24/2024
Operations
Publications
Riquier S, Bessiere C, Guibert B, Bouge A, Boureux A, Ruffle F, Audoux J, Gilbert N, Xue H, Gautheret D, Commes T. Kmerator Suite: design of specific <i>k</i> -mer signatures and automatic metadata discovery in large RNA-seq datasets. NAR Genomics and Bioinformatics. 2021;3(3). doi:10.1093/nargab/lqab058. PMID:34179780. PMCID:PMC8221386.