pypgatk
pypgatk: Python Library for Proteogenomics Database Construction and Variant Peptide Discovery
pypgatk constructs proteogenomics databases and enables identification of novel and variant protein-coding transcripts by integrating ENSEMBL resources and performing multi-frame translation of non-canonical transcripts.
Key Features:
- Three-Frame Translation: Performs three-frame translation of pseudogenes, long non-coding RNAs (lncRNAs), and non-canonical transcripts derived from alternative splicing events to identify novel protein-coding sequences.
- Exonic Out-of-Frame Translation: Generates alternative protein sequences from canonical protein-coding mRNAs through exonic out-of-frame translation.
- Variant Integration: Incorporates genomic variants from COSMIC, cBioPortal, gnomAD, and patient-derived sequencing data to produce customized proteogenomics databases.
- Target/Decoy Generation: Implements optimized target/decoy database construction using the DecoyPyrat algorithm to support proteomics data analysis.
- pgdb Workflow Integration: Interfaces with the pgdb workflow to automate ENSEMBL-based proteogenomics database generation.
Scientific Applications:
- Non-Canonical Peptide Identification: Enables detection of cryptic and non-canonical peptides in mass spectrometry datasets, including large-scale reanalysis of PRIDE datasets and cell-type specific database generation.
- Proteogenomics Research: Supports systematic identification and characterization of novel protein-coding transcripts and variant-derived peptides.
Methodology:
pypgatk integrates ENSEMBL annotations with genomic variant data to generate customized proteogenomics sequence databases. It applies three-frame and exonic out-of-frame translation strategies to canonical and non-canonical transcripts, followed by optimized target/decoy construction using the DecoyPyrat algorithm to enable downstream mass spectrometry-based peptide identification.
Topics
Details
- License:
- Apache-2.0
- Cost:
- Free of charge
- Tool Type:
- library
- Operating Systems:
- Mac, Linux, Windows
- Programming Languages:
- Python
- Added:
- 10/24/2021
- Last Updated:
- 10/24/2021
Operations
Publications
Umer HM, Zhu Y, Pfeuffer J, Sachsenberg T, Lehtiö J, Branca R, Perez-Riverol Y. Generation of ENSEMBL-based proteogenomics databases boosts the identification of non-canonical peptides. Unknown Journal. 2021. doi:10.1101/2021.06.08.447496.