Psafe

Psafe corrects biases and errors in pooled next-generation sequencing (NGS) data to improve allele frequency estimation for studies of common and rare single nucleotide polymorphisms (SNPs).


Key Features:

  • Bias correction: Adjusts allele frequency estimates to account for biases and errors identified in pooled sequencing data.
  • Use of targeted sequencing and microarray genotypes: Leverages targeted pooled sequencing datasets together with microarray SNP genotypes from the same subjects to characterize bias sources.
  • PCR bias identification: Identifies and quantifies PCR amplification biases occurring during target capture and sequencing.
  • Reference allele preferential bias detection: Detects and quantifies reference allele preferential bias affecting allele calls and frequency estimates.
  • Simulation-based evaluation: Uses extensive simulations to demonstrate impacts of identified biases on allele frequency accuracy and association test power.
  • Workflow protocol for adjustment: Implements a workflow protocol that adjusts pooled sequencing data for the quantified biases and errors.
  • Pooled sequencing context awareness: Accounts for deviations from simplified pooling assumptions such as non-uniform sample contributions and variable sequencing depth.

Scientific Applications:

  • Allele frequency estimation: Improves accuracy of allele frequency estimates in pooled NGS studies.
  • Association testing: Enhances statistical power and reliability of association tests in pooled sequencing experiments.
  • Variant discovery in NGS: Supports analysis of common and rare SNPs from pooled sequencing without prior SNP knowledge.
  • Bias assessment and mitigation: Enables quantification and correction of PCR amplification and reference allele preferential biases in pooled datasets.

Methodology:

Psafe analyzes targeted pooled sequencing datasets alongside microarray SNP genotypes to identify and quantify biases (including PCR amplification during target capture and sequencing and reference allele preferential bias), uses extensive simulations to evaluate their impact on allele frequency estimation and association test power, and applies a workflow protocol to adjust for those biases.

Topics

Details

Tool Type:
command-line tool
Operating Systems:
Linux, Windows, Mac
Programming Languages:
R, Perl
Added:
8/3/2017
Last Updated:
11/25/2024

Operations

Publications

Chen X, Listman JB, Slack FJ, Gelernter J, Zhao H. Biases and Errors on Allele Frequency Estimation and Disease Association Tests of Next‐Generation Sequencing of Pooled Samples. Genetic Epidemiology. 2012;36(6):549-560. doi:10.1002/gepi.21648. PMID:22674656. PMCID:PMC3477622.

PMID: 22674656
PMCID: PMC3477622
Funding: - NIH: RR19895 - National Institutes of Health: N01-HG-65403

Documentation

Links