pypgatk

pypgatk: Python Library for Proteogenomics Database Construction and Variant Peptide Discovery

pypgatk constructs proteogenomics databases and enables identification of novel and variant protein-coding transcripts by integrating ENSEMBL resources and performing multi-frame translation of non-canonical transcripts.


Key Features:

  • Three-Frame Translation: Performs three-frame translation of pseudogenes, long non-coding RNAs (lncRNAs), and non-canonical transcripts derived from alternative splicing events to identify novel protein-coding sequences.
  • Exonic Out-of-Frame Translation: Generates alternative protein sequences from canonical protein-coding mRNAs through exonic out-of-frame translation.
  • Variant Integration: Incorporates genomic variants from COSMIC, cBioPortal, gnomAD, and patient-derived sequencing data to produce customized proteogenomics databases.
  • Target/Decoy Generation: Implements optimized target/decoy database construction using the DecoyPyrat algorithm to support proteomics data analysis.
  • pgdb Workflow Integration: Interfaces with the pgdb workflow to automate ENSEMBL-based proteogenomics database generation.

Scientific Applications:

  • Non-Canonical Peptide Identification: Enables detection of cryptic and non-canonical peptides in mass spectrometry datasets, including large-scale reanalysis of PRIDE datasets and cell-type specific database generation.
  • Proteogenomics Research: Supports systematic identification and characterization of novel protein-coding transcripts and variant-derived peptides.

Methodology:

pypgatk integrates ENSEMBL annotations with genomic variant data to generate customized proteogenomics sequence databases. It applies three-frame and exonic out-of-frame translation strategies to canonical and non-canonical transcripts, followed by optimized target/decoy construction using the DecoyPyrat algorithm to enable downstream mass spectrometry-based peptide identification.

Topics

Details

License:
Apache-2.0
Cost:
Free of charge
Tool Type:
library
Operating Systems:
Mac, Linux, Windows
Programming Languages:
Python
Added:
10/24/2021
Last Updated:
10/24/2021

Operations

Publications

Umer HM, Zhu Y, Pfeuffer J, Sachsenberg T, Lehtiö J, Branca R, Perez-Riverol Y. Generation of ENSEMBL-based proteogenomics databases boosts the identification of non-canonical peptides. Unknown Journal. 2021. doi:10.1101/2021.06.08.447496.

Documentation

Links