DupChecker
DupChecker identifies duplicated samples in high-throughput genomic datasets by generating MD5 fingerprints to ensure data integrity for meta-analyses of gene expression data.
Key Features:
- MD5 fingerprinting: Generates MD5 fingerprints from raw data to detect identical samples.
- Efficient identification: Detects duplicates efficiently without requiring extensive computational resources.
- Application in gene expression data: Targets gene expression datasets where duplication can bias meta-analyses, clustering, and model fitting.
- Pre-analysis screening: Intended to be applied prior to downstream analyses to prevent duplicate-induced false positives and overfitting.
- Bioconductor package: Distributed as a Bioconductor package for integration into analysis workflows.
Scientific Applications:
- Meta-analysis quality control: Used in meta-analyses of high-throughput genomic data to detect and remove duplicated samples.
- Reduction of analytical artifacts: Reduces false positives, misleading clustering, and model overfitting caused by duplicated samples.
- Data integrity verification: Provides MD5-based fingerprints to verify sample identity across public genomic databases.
Methodology:
Generate unique MD5 hashes for each sample from raw data and compare these fingerprints across datasets to identify duplicate samples.
Topics
Collections
Details
- License:
- GPL-2.0
- Tool Type:
- command-line tool, library
- Operating Systems:
- Linux, Windows, Mac
- Programming Languages:
- R
- Added:
- 1/17/2017
- Last Updated:
- 1/10/2019
Operations
Publications
Sheng Q, Shyr Y, Chen X. DupChecker: a bioconductor package for checking high-throughput genomic data redundancy in meta-analysis. BMC Bioinformatics. 2014;15(1). doi:10.1186/1471-2105-15-323. PMID:25267467. PMCID:PMC4261523.