DeconSeq
DeconSeq identifies and removes contaminant DNA sequences from metagenomic datasets generated by high-throughput sequencing to improve accuracy of downstream genome assembly and microbial diversity analyses.
Key Features:
- Automated detection and removal: Rapidly identifies and removes contaminant sequences from longer-read datasets (mean read length of 150 bp).
- Host contamination screening: Detects contamination from human and other host sequences in metagenomic data.
- Alignment technology: Uses a modified version of the BWA-SW aligner to process longer reads for contaminant detection.
- Classification of contaminants: Categorizes potential contaminant sequences and classifies alignment results.
- Redundant-hit elimination: Removes redundant hits by selecting higher similarity to non-contaminant genomes.
- Visualization of results: Produces graphical visualizations of alignment results and contaminant classifications.
- Empirical detection in published datasets: Identified possible human DNA contamination in 145 (72%) of 202 previously published microbial and viral metagenomes, with some datasets containing up to 64% contaminating sequences.
- Quality assurance indicators: Reveals whether a sequencing experiment succeeded, whether the correct sample was sequenced, and whether contamination from DNA preparation or host is present.
Scientific Applications:
- Metagenomic quality control: Screening and removal of contaminant sequences to improve data quality for downstream analyses.
- Genome assembly improvement: Removing contaminant reads to reduce misassembly of sequence contigs.
- Microbial diversity studies: Ensuring accurate assessment of microbial community composition by eliminating non-target DNA.
- Virology and microbiology research: Detecting host-derived contamination in viral and microbial metagenomes.
- Retrospective dataset screening: Screening published metagenomic datasets for human or host contamination.
Methodology:
Aligns reads using a modified BWA-SW, automatically identifies and removes contaminant sequences from longer-read datasets (mean read length of 150 bp), categorizes potential contaminants, eliminates redundant hits by choosing higher similarity to non-contaminant genomes, and generates graphical visualizations of alignments and classifications.
Topics
Collections
Details
- License:
- GPL-3.0
- Maturity:
- Mature
- Tool Type:
- api
- Operating Systems:
- Linux, Mac
- Programming Languages:
- Perl, C
- Added:
- 1/13/2017
- Last Updated:
- 11/25/2024
Operations
Publications
Schmieder R, Edwards R. Fast Identification and Removal of Sequence Contamination from Genomic and Metagenomic Datasets. PLoS ONE. 2011;6(3):e17288. doi:10.1371/journal.pone.0017288. PMID:21408061. PMCID:PMC3052304.