Kontaminant

Kontaminant performs k-mer-based screening and host-sequence subtraction to reduce contaminant reads and enrich putative viral nucleic acids from next-generation sequencing (NGS) data derived from eukaryotic tissue samples.


Key Features:

  • K-mer-Based Filtering: Employs a k-mer frequency-based filter to reduce the number of contigs prior to assembly and focus on putative viral nucleic acids.
  • Pre-Assembly Host-Mapping Subtraction: Integrates short-read mapping tools to subtract host sequences before assembly.
  • Low Complexity Filter: Removes low-complexity, non-informative reads to refine input datasets for assembly.

Scientific Applications:

  • Viral discovery from eukaryotic tissues: Enables identification of novel viral sequences in complex eukaryotic tissue NGS datasets.
  • Contig reduction and assembly improvement: Reduces the number of assembled contigs (reported up to 99.97%) and increases viral contig size to simplify downstream analysis.
  • Validation and benchmarking: Has been validated on Illumina sequencing of naturally infected liver tissue and on simulated datasets.

Methodology:

Applies k-mer frequency-based filtering, low-complexity read filtering, and short-read mapping-based host subtraction prior to assembly.

Topics

Details

Tool Type:
command-line tool
Operating Systems:
Linux
Programming Languages:
C
Added:
12/18/2017
Last Updated:
12/10/2018

Operations

Publications

Leggett RM, Ramirez-Gonzalez RH, Clavijo BJ, Waite D, Davey RP. Sequencing quality assessment tools to enable data-driven informatics for high throughput genomics. Frontiers in Genetics. 2013;4. doi:10.3389/fgene.2013.00288. PMID:24381581. PMCID:PMC3865868.

Daly GM, Leggett RM, Rowe W, Stubbs S, Wilkinson M, Ramirez-Gonzalez RH, Caccamo M, Bernal W, Heeney JL. Host Subtraction, Filtering and Assembly Validations for Novel Viral Discovery Using Next Generation Sequencing Data. PLOS ONE. 2015;10(6):e0129059. doi:10.1371/journal.pone.0129059. PMID:26098299. PMCID:PMC4476701.

Documentation

Links