PrecisionProDB
PrecisionProDB generates personalized protein databases from genomic data to improve peptide search accuracy in mass spectrometry and support precision medicine proteogenomics.
Key Features:
- Personalized database generation: Generates customized protein databases tailored to an individual's genomic profile for peptide search applications in mass spectrometry.
- Python-based implementation: Implemented in Python for computational generation of personalized protein databases.
- Support for multiple formats and references: Supports various popular file formats and reference databases to ensure compatibility with existing workflows.
- Rapid database generation: Produces personalized protein databases in a matter of minutes.
- Population-specific reference databases: Enables creation of human population-specific protein databases that demonstrated an average 0.34% improvement in peptide identification rates.
- Cell line-specific variant incorporation: Incorporates cell line-specific genetic variants and achieved a 0.71% enhancement in peptide identification for the Jurkat cell line.
- Genomics–proteomics integration: Integrates next-generation sequencing (NGS) genomic data with proteomic peptide searches to improve detection power and accuracy.
Scientific Applications:
- Proteogenomics studies: Improves protein identification accuracy in proteogenomics analyses by incorporating sample-specific genomic variants into search databases.
- Precision medicine: Enhances detection of individual-specific proteomic variants relevant to precision medicine investigations.
- Population-level proteomics: Supports generation of population-specific reference protein sets to assess population-linked peptide identification differences.
- Cell line proteomics: Facilitates incorporation of cell line-specific variants, exemplified by improved peptide identification in the Jurkat cell line.
- NGS-integrated workflows: Enables workflows that combine next-generation sequencing data with mass spectrometry peptide searches to increase search sensitivity and accuracy.
Methodology:
Implemented in Python to generate personalized protein databases from genomic variant data for peptide searches in mass spectrometry, integrating NGS-derived variant information into reference protein sequences.
Topics
Details
- License:
- GPL-3.0
- Tool Type:
- library, workflow
- Programming Languages:
- Python
- Added:
- 11/22/2021
- Last Updated:
- 11/22/2021
Operations
Publications
Cao X, Xing J. PrecisionProDB: improving the proteomics performance for precision medicine. Bioinformatics. 2021;37(19):3361-3363. doi:10.1093/bioinformatics/btab218. PMID:33787868.
PMID: 33787868
Documentation
Citation instructions
https://github.com/ATPs/PrecisionProDB_referencesLinks
Issue tracker
https://github.com/ATPs/PrecisionProDB/issues