decontam

decontam identifies contaminant DNA sequences in marker-gene and metagenomic sequencing (MGS) data to improve the accuracy of microbial community surveys.


Key Features:

  • Statistical classification: Implements a statistical classification procedure to distinguish contaminants from true biological features in sequencing data.
  • Frequency–concentration pattern: Uses the tendency of contaminants to appear at higher frequency in samples with lower DNA concentrations as a discriminating signal.
  • Negative-control prevalence: Uses presence and prevalence of features in negative control samples as a criterion for identifying contaminants.
  • Applicability to environmental sequencing features: Applies to any feature derived from environmental sequencing data, including marker-gene and metagenomic-derived features (e.g., OTUs or ASVs).
  • Detection of reagent-derived contaminants: Targets contaminant DNA sequences that can originate from reagents and other laboratory sources.

Scientific Applications:

  • Microbial community analysis: Improves accuracy of microbial community profiles by identifying and removing contaminant sequences from marker-gene and MGS datasets, including human oral microbiome studies.
  • Reduction of technical variation: Reduces technical variation in datasets such as dilution series and those generated by different sequencing protocols by removing contaminant-driven signals.
  • Validation of biological findings: Helps validate published findings by clarifying cases where low-frequency taxa represent contaminants, including analyses related to the placenta microbiome and taxa previously associated with preterm birth.

Methodology:

Performs statistical classification that leverages two explicit patterns—higher contaminant frequency in low-DNA-concentration samples and occurrence in negative controls—to distinguish contaminant sequences from true microbial features.

Topics

Collections

Details

License:
Artistic-2.0
Cost:
Free of charge
Tool Type:
library
Operating Systems:
Linux, Windows, Mac
Programming Languages:
R
Added:
7/17/2018
Last Updated:
12/10/2018

Operations

Publications

Davis NM, Proctor DM, Holmes SP, Relman DA, Callahan BJ. Simple statistical identification and removal of contaminant sequences in marker-gene and metagenomics data. Unknown Journal. 2017. doi:10.1101/221499.

Documentation

Downloads

Links