Kmerator

Kmerator predicts and analyzes specific k-mers from RNA-sequencing (RNA-seq) data to enable k-mer based transcript quantification and signature extraction using a reference genome and an ENSEMBL-like transcriptome.


Key Features:

  • Specific k-mer prediction: Predicts and analyzes specific k-mers (tags) derived from input sequences for gene- and transcript-level signatures.
  • Contig generation: Generates contigs as sequences of consecutive k-mers overlapping by k-1 from specific k-mers.
  • Jellyfish indexing: Uses Jellyfish to build two requestable indexes from the reference genome and the ENSEMBL-like transcriptome.
  • K-mer decomposition and counting: Decomposes input transcript or gene sequences into k-mers and counts occurrences within both the genome and the transcriptome.
  • Human gene coverage and accuracy: Produces specific k-mers for 97% of human genes and enables accurate gene expression quantification in simulated datasets.
  • Suite composition: Forms part of the Kmerator Suite comprising three tools that extract, quantify, and analyze k-mer signatures, including KmerExploR.
  • Metadata inference (KmerExploR): Applies gene-specific k-mer predictors to infer metadata such as library protocol, sample features, and potential contaminations.
  • Advanced queries: Supports queries for known or novel biomarkers, including identification of mutations, gene fusions, and long non-coding RNAs.
  • Computational efficiency: Provides a k-mer decomposition alternative to alignment-based methods to reduce computational resource requirements for large-scale RNA-seq analyses.

Scientific Applications:

  • Gene expression quantification: Quantifies transcript- and gene-level expression from RNA-seq via specific k-mer counts.
  • Sample metadata inference: Infers library protocol, sample features, and contaminations from RNA-seq datasets using KmerExploR.
  • Biomarker discovery: Enables queries for known or novel biomarkers from k-mer signatures.
  • Detection of genetic events: Identifies mutations and gene fusions from k-mer patterns.
  • Long non-coding RNA analysis: Detects long non-coding RNAs using k-mer based signatures.
  • Large-scale RNA-seq processing: Facilitates resource-efficient analysis of large RNA-seq cohorts using k-mer decomposition instead of alignment.

Methodology:

Uses Jellyfish to create two requestable indexes from the reference genome and the ENSEMBL-like transcriptome; decomposes input transcript or gene sequences into k-mers and counts their occurrences within genome and transcriptome; generates contigs by joining consecutive specific k-mers overlapping by k-1; predicts specific k-mers (tags) derived from input sequences.

Topics

Details

License:
GPL-3.0
Maturity:
Emerging
Cost:
Free of charge
Tool Type:
command-line tool
Programming Languages:
Python
Added:
10/4/2021
Last Updated:
11/24/2024

Operations

Publications

Riquier S, Bessiere C, Guibert B, Bouge A, Boureux A, Ruffle F, Audoux J, Gilbert N, Xue H, Gautheret D, Commes T. Kmerator Suite: design of specific <i>k</i> -mer signatures and automatic metadata discovery in large RNA-seq datasets. NAR Genomics and Bioinformatics. 2021;3(3). doi:10.1093/nargab/lqab058. PMID:34179780. PMCID:PMC8221386.

PMID: 34179780
PMCID: PMC8221386
Funding: - Agence Nationale de la Recherche: ANR-10-INBS-09 - Canceropole Grand Ouest: 2017-EM24 - Region Occitanie: R19073FF

Links