DeconSeq

DeconSeq identifies and removes contaminant DNA sequences from metagenomic datasets generated by high-throughput sequencing to improve accuracy of downstream genome assembly and microbial diversity analyses.


Key Features:

  • Automated detection and removal: Rapidly identifies and removes contaminant sequences from longer-read datasets (mean read length of 150 bp).
  • Host contamination screening: Detects contamination from human and other host sequences in metagenomic data.
  • Alignment technology: Uses a modified version of the BWA-SW aligner to process longer reads for contaminant detection.
  • Classification of contaminants: Categorizes potential contaminant sequences and classifies alignment results.
  • Redundant-hit elimination: Removes redundant hits by selecting higher similarity to non-contaminant genomes.
  • Visualization of results: Produces graphical visualizations of alignment results and contaminant classifications.
  • Empirical detection in published datasets: Identified possible human DNA contamination in 145 (72%) of 202 previously published microbial and viral metagenomes, with some datasets containing up to 64% contaminating sequences.
  • Quality assurance indicators: Reveals whether a sequencing experiment succeeded, whether the correct sample was sequenced, and whether contamination from DNA preparation or host is present.

Scientific Applications:

  • Metagenomic quality control: Screening and removal of contaminant sequences to improve data quality for downstream analyses.
  • Genome assembly improvement: Removing contaminant reads to reduce misassembly of sequence contigs.
  • Microbial diversity studies: Ensuring accurate assessment of microbial community composition by eliminating non-target DNA.
  • Virology and microbiology research: Detecting host-derived contamination in viral and microbial metagenomes.
  • Retrospective dataset screening: Screening published metagenomic datasets for human or host contamination.

Methodology:

Aligns reads using a modified BWA-SW, automatically identifies and removes contaminant sequences from longer-read datasets (mean read length of 150 bp), categorizes potential contaminants, eliminates redundant hits by choosing higher similarity to non-contaminant genomes, and generates graphical visualizations of alignments and classifications.

Topics

Collections

Details

License:
GPL-3.0
Maturity:
Mature
Tool Type:
api
Operating Systems:
Linux, Mac
Programming Languages:
Perl, C
Added:
1/13/2017
Last Updated:
11/25/2024

Operations

Publications

Schmieder R, Edwards R. Fast Identification and Removal of Sequence Contamination from Genomic and Metagenomic Datasets. PLoS ONE. 2011;6(3):e17288. doi:10.1371/journal.pone.0017288. PMID:21408061. PMCID:PMC3052304.