decontam
decontam identifies contaminant DNA sequences in marker-gene and metagenomic sequencing (MGS) data to improve the accuracy of microbial community surveys.
Key Features:
- Statistical classification: Implements a statistical classification procedure to distinguish contaminants from true biological features in sequencing data.
- Frequency–concentration pattern: Uses the tendency of contaminants to appear at higher frequency in samples with lower DNA concentrations as a discriminating signal.
- Negative-control prevalence: Uses presence and prevalence of features in negative control samples as a criterion for identifying contaminants.
- Applicability to environmental sequencing features: Applies to any feature derived from environmental sequencing data, including marker-gene and metagenomic-derived features (e.g., OTUs or ASVs).
- Detection of reagent-derived contaminants: Targets contaminant DNA sequences that can originate from reagents and other laboratory sources.
Scientific Applications:
- Microbial community analysis: Improves accuracy of microbial community profiles by identifying and removing contaminant sequences from marker-gene and MGS datasets, including human oral microbiome studies.
- Reduction of technical variation: Reduces technical variation in datasets such as dilution series and those generated by different sequencing protocols by removing contaminant-driven signals.
- Validation of biological findings: Helps validate published findings by clarifying cases where low-frequency taxa represent contaminants, including analyses related to the placenta microbiome and taxa previously associated with preterm birth.
Methodology:
Performs statistical classification that leverages two explicit patterns—higher contaminant frequency in low-DNA-concentration samples and occurrence in negative controls—to distinguish contaminant sequences from true microbial features.
Topics
Collections
Details
- License:
- Artistic-2.0
- Cost:
- Free of charge
- Tool Type:
- library
- Operating Systems:
- Linux, Windows, Mac
- Programming Languages:
- R
- Added:
- 7/17/2018
- Last Updated:
- 12/10/2018
Operations
Publications
Davis NM, Proctor DM, Holmes SP, Relman DA, Callahan BJ. Simple statistical identification and removal of contaminant sequences in marker-gene and metagenomics data. Unknown Journal. 2017. doi:10.1101/221499.
DOI: 10.1101/221499