MSPolygraph

MSPolygraph identifies peptides from tandem mass spectrometry (MS/MS) data by performing hybrid database and spectral library searches combined with statistical models and machine learning to improve peptide-spectrum matching.


Key Features:

  • Hybrid database and spectral library search: Performs searches against sequence databases and spectral libraries for MS/MS peptide identification.
  • Statistical hypothesis testing: Evaluates two-hypothesis tests (H(0): match by chance; H(A): spectrum produced by a specific peptide) for peptide-spectrum matches.
  • Probability models for fragment ions: Uses initial models with uniform probabilities across labile bonds and fragment ion types and advanced models with non-uniform probability distributions for each ion type under both hypotheses.
  • Model comparison and evaluation: Compares statistical models by assessing identification rates on independent datasets and evaluates standard likelihood models against information-theory approaches derived from those likelihoods.
  • Peak intensity integration: Incorporates peak intensities into scoring functions to enhance the reliability of peptide-spectrum matches.
  • Machine learning (SVM): Employs a support-vector machine using multiple characteristics of the scoring functions to distinguish correct from incorrect identifications and reduce misidentification rates relative to cross-correlation methods.

Scientific Applications:

  • Peptide Identification: Improves accuracy of peptide identification from MS/MS spectra through refined statistical scoring and intensity-aware models.
  • Proteomics Research: Supports protein characterization and quantification by providing higher-confidence peptide-spectrum matches for proteomic analyses.
  • Data Analysis Optimization: Enables assessment and optimization of scoring models and machine learning classifiers for complex mass spectrometry datasets.

Methodology:

Performs two-hypothesis testing with probability models for fragment ions (uniform and non-uniform), compares models by identification rates on independent datasets, derives and evaluates likelihood and information-theory-based scores, integrates peak intensities into scoring functions, and trains a support-vector machine on scoring-function characteristics to classify identifications.

Topics

Collections

Details

Tool Type:
command-line tool
Operating Systems:
Linux
Programming Languages:
C
Added:
1/17/2017
Last Updated:
11/25/2024

Operations

Data Inputs & Outputs

Publications

Cannon WR, Jarman KH, Webb-Robertson BM, Baxter DJ, Oehmen CS, Jarman KD, Heredia-Langner A, Auberry KJ, Anderson GA. Comparison of Probability and Likelihood Models for Peptide Identification from Tandem Mass Spectrometry Data. Journal of Proteome Research. 2005;4(5):1687-1698. doi:10.1021/pr050147v. PMID:16212422.

Documentation

Downloads

Links

Software catalogue
http://ms-utils.org