MSPolygraph
MSPolygraph identifies peptides from tandem mass spectrometry (MS/MS) data by performing hybrid database and spectral library searches combined with statistical models and machine learning to improve peptide-spectrum matching.
Key Features:
- Hybrid database and spectral library search: Performs searches against sequence databases and spectral libraries for MS/MS peptide identification.
- Statistical hypothesis testing: Evaluates two-hypothesis tests (H(0): match by chance; H(A): spectrum produced by a specific peptide) for peptide-spectrum matches.
- Probability models for fragment ions: Uses initial models with uniform probabilities across labile bonds and fragment ion types and advanced models with non-uniform probability distributions for each ion type under both hypotheses.
- Model comparison and evaluation: Compares statistical models by assessing identification rates on independent datasets and evaluates standard likelihood models against information-theory approaches derived from those likelihoods.
- Peak intensity integration: Incorporates peak intensities into scoring functions to enhance the reliability of peptide-spectrum matches.
- Machine learning (SVM): Employs a support-vector machine using multiple characteristics of the scoring functions to distinguish correct from incorrect identifications and reduce misidentification rates relative to cross-correlation methods.
Scientific Applications:
- Peptide Identification: Improves accuracy of peptide identification from MS/MS spectra through refined statistical scoring and intensity-aware models.
- Proteomics Research: Supports protein characterization and quantification by providing higher-confidence peptide-spectrum matches for proteomic analyses.
- Data Analysis Optimization: Enables assessment and optimization of scoring models and machine learning classifiers for complex mass spectrometry datasets.
Methodology:
Performs two-hypothesis testing with probability models for fragment ions (uniform and non-uniform), compares models by identification rates on independent datasets, derives and evaluates likelihood and information-theory-based scores, integrates peak intensities into scoring functions, and trains a support-vector machine on scoring-function characteristics to classify identifications.
Topics
Collections
Details
- Tool Type:
- command-line tool
- Operating Systems:
- Linux
- Programming Languages:
- C
- Added:
- 1/17/2017
- Last Updated:
- 11/25/2024
Operations
Data Inputs & Outputs
Blind peptide database search
Publications
Cannon WR, Jarman KH, Webb-Robertson BM, Baxter DJ, Oehmen CS, Jarman KD, Heredia-Langner A, Auberry KJ, Anderson GA. Comparison of Probability and Likelihood Models for Peptide Identification from Tandem Mass Spectrometry Data. Journal of Proteome Research. 2005;4(5):1687-1698. doi:10.1021/pr050147v. PMID:16212422.