doppelgangR

doppelgangR identifies duplicate samples within genomic expression datasets to detect unrecognized sample duplication that can inflate predictive accuracy and confound differential expression analyses, with particular focus on cancer transcriptomes.


Key Features:

  • Minimal Input Requirement: Operates with minimal input and primarily requires lists of ExpressionSet objects.
  • Comprehensive Search Mechanism: Conducts exhaustive pairwise comparisons across provided lists to detect duplicated samples using genomic expression data, phenotype/clinical data and unique identifiers in pData(eset), often referred to as "smoking guns."
  • Cross-Technology Compatibility: Matches duplicate transcriptomes across different microarray technologies and between microarray and RNA sequencing even when nucleotide-level sequence data are unavailable.

Scientific Applications:

  • Data Integrity Checking: Provides routine screening to identify duplicate samples in cancer transcriptome datasets to improve dataset integrity.
  • Database Screening: Enables screening of transcriptome databases to detect duplicated or highly similar profiles prior to reuse or meta-analysis.
  • Improving Analysis Reliability: Reduces the risk of false inflation of prediction accuracy and overstated confidence in differential expression studies by identifying duplicates.
  • Workflow Integration: Supports integration into standard procedures for combining multiple genomic datasets to ensure more reliable downstream analyses.

Methodology:

The package analyzes whole-genome expression profiles provided as ExpressionSet objects and performs exhaustive pairwise comparisons using genomic expression data, phenotype/clinical pData fields and unique identifiers in pData(eset) to detect probable duplicate samples across microarray platforms and RNA-seq datasets; it has been demonstrated on ovarian, breast, bladder and colorectal microarray profiles and matched TCGA profiles.

Topics

Collections

Details

License:
GPL-2.0
Tool Type:
command-line tool, library
Operating Systems:
Linux, Windows, Mac
Programming Languages:
R
Added:
1/17/2017
Last Updated:
11/25/2024

Operations

Publications

Waldron L, Riester M, Ramos M, Parmigiani G, Birrer M. The Doppelgänger Effect: Hidden Duplicates in Databases of Transcriptome Profiles. Journal of the National Cancer Institute. 2016;108(11):djw146. doi:10.1093/jnci/djw146. PMID:27381624. PMCID:PMC5241903.

Documentation

Downloads

Links