PrecisionProDB

PrecisionProDB generates personalized protein databases from genomic data to improve peptide search accuracy in mass spectrometry and support precision medicine proteogenomics.


Key Features:

  • Personalized database generation: Generates customized protein databases tailored to an individual's genomic profile for peptide search applications in mass spectrometry.
  • Python-based implementation: Implemented in Python for computational generation of personalized protein databases.
  • Support for multiple formats and references: Supports various popular file formats and reference databases to ensure compatibility with existing workflows.
  • Rapid database generation: Produces personalized protein databases in a matter of minutes.
  • Population-specific reference databases: Enables creation of human population-specific protein databases that demonstrated an average 0.34% improvement in peptide identification rates.
  • Cell line-specific variant incorporation: Incorporates cell line-specific genetic variants and achieved a 0.71% enhancement in peptide identification for the Jurkat cell line.
  • Genomics–proteomics integration: Integrates next-generation sequencing (NGS) genomic data with proteomic peptide searches to improve detection power and accuracy.

Scientific Applications:

  • Proteogenomics studies: Improves protein identification accuracy in proteogenomics analyses by incorporating sample-specific genomic variants into search databases.
  • Precision medicine: Enhances detection of individual-specific proteomic variants relevant to precision medicine investigations.
  • Population-level proteomics: Supports generation of population-specific reference protein sets to assess population-linked peptide identification differences.
  • Cell line proteomics: Facilitates incorporation of cell line-specific variants, exemplified by improved peptide identification in the Jurkat cell line.
  • NGS-integrated workflows: Enables workflows that combine next-generation sequencing data with mass spectrometry peptide searches to increase search sensitivity and accuracy.

Methodology:

Implemented in Python to generate personalized protein databases from genomic variant data for peptide searches in mass spectrometry, integrating NGS-derived variant information into reference protein sequences.

Topics

Details

License:
GPL-3.0
Tool Type:
library, workflow
Programming Languages:
Python
Added:
11/22/2021
Last Updated:
11/22/2021

Operations

Publications

Cao X, Xing J. PrecisionProDB: improving the proteomics performance for precision medicine. Bioinformatics. 2021;37(19):3361-3363. doi:10.1093/bioinformatics/btab218. PMID:33787868.

Documentation

Links