PGA
PGA generates customized protein databases from RNA-Seq and enhances peptide identification from tandem mass spectrometry to detect novel peptides absent from reference protein databases.
Key Features:
- Customized Protein Database Construction: Constructs protein databases from RNA-Seq using reference-guided or reference-independent modes and incorporates reference sequences from Refseq and ENSEMBL alongside novel peptide sequences.
- R package implementation: Implemented as an R package that automates processing of tandem mass spectrometry (MS/MS) data obtained from various MS platforms.
- Novel Peptide Identification: Enables detection of single amino acid polymorphism (SAP) peptides, INDEL peptides, splice-junction peptides, and novel transcript-derived peptides not present in standard reference databases.
Scientific Applications:
- Proteomics discovery: Expands peptide detection beyond reference databases to discover novel peptides in proteomics studies.
- Genetic variation and splicing analysis: Supports identification of peptides arising from SAPs, INDELs, and alternative splicing events.
- Disease mechanisms and biomarker/target discovery: Facilitates identification of novel peptide-derived biomarkers and candidate therapeutic targets related to protein function and disease mechanisms.
Methodology:
Implemented as an R package that constructs customized protein databases from RNA-Seq (reference-guided or independent), incorporates Refseq and ENSEMBL sequences, and processes tandem mass spectrometry (MS/MS) data from various MS platforms to identify SAP, INDEL, splice-junction and novel transcript-derived peptides.
Topics
Collections
Details
- License:
- GPL-2.0
- Tool Type:
- command-line tool, library
- Operating Systems:
- Linux, Windows, Mac
- Programming Languages:
- R
- Added:
- 1/17/2017
- Last Updated:
- 7/20/2019
Operations
Publications
Wen B, Xu S, Zhou R, Zhang B, Wang X, Liu X, Xu X, Liu S. PGA: an R/Bioconductor package for identification of novel peptides using a customized database derived from RNA-Seq. BMC Bioinformatics. 2016;17(1). doi:10.1186/s12859-016-1133-3. PMID:27316337. PMCID:PMC4912784.