ContEst
ContEst estimates cross-individual contamination levels in next-generation sequencing (NGS) datasets as part of the Genome Analysis Toolkit (GATK) suite to support quality control in genomic and cancer sequencing analyses.
Key Features:
- Accurate Estimation: Estimates contamination across varying sources, read depths, and known concentrations by detecting and quantifying sequences that do not belong to the intended sample.
- Validation with in silico mixes: Validated using sequencing data mixed in silico at predetermined concentrations to assess accuracy across diverse experimental conditions.
- Application to published cancer datasets: Applied to published cancer sequencing datasets to measure contamination in real-world studies.
- GATK integration: Implemented as a module within the Genome Analysis Toolkit (GATK), enabling integration with GATK computational workflows.
- Contamination metrics: Produces actionable contamination metrics for downstream quality-control and analysis.
Scientific Applications:
- Genomic and cancer research quality control: Assess and ensure the validity of genomic analyses in cancer research and other fields by quantifying cross-individual contamination.
- Bias adjustment: Enable adjustment for potential biases introduced by cross-individual DNA to improve the robustness of findings.
- Large-scale genomics projects: Enhance quality-control processes in large-scale sequencing and genomics projects.
Methodology:
Operates as a module within the Genome Analysis Toolkit (GATK), analyzing sequencing data to detect and quantify non-sample sequences; validation was performed using in silico mixed sequencing data at predetermined concentrations.
Topics
Details
- License:
- BSD-3-Clause
- Tool Type:
- command-line tool
- Operating Systems:
- Linux, Mac
- Programming Languages:
- Java
- Added:
- 1/13/2017
- Last Updated:
- 11/25/2024
Operations
Publications
Cibulskis K, McKenna A, Fennell T, Banks E, DePristo M, Getz G. ContEst: estimating cross-contamination of human samples in next-generation sequencing data. Bioinformatics. 2011;27(18):2601-2602. doi:10.1093/bioinformatics/btr446. PMID:21803805. PMCID:PMC3167057.