STEPP

STEPP predicts proteotypic peptides for mass spectrometry-based proteomics using a support vector machine to improve peptide identification.


Key Features:

  • Quantitative Prediction: Quantitatively predicts proteotypic peptides to prioritize peptides most likely to be detected by mass spectrometry.
  • Support Vector Machine model: Employs a support vector machine (SVM) that uses a descriptor space derived from 35 amino-acid properties including content, charge, hydrophilicity, and polarity.
  • Validation on AMT databases: Validated on three independently derived accurate mass and elution time (AMT) databases from Shewanella oneidensis, Salmonella typhimurium, and Yersinia pestis, achieving an average accuracy of approximately 0.83 with standard deviation < 0.038.
  • Reduced descriptor set: Maintains high prediction accuracy when the descriptor space is reduced to a small set of 13 variables.
  • Cross-species validation: Demonstrates robustness across different species used in validation, supporting applicability in diverse proteomic datasets.

Scientific Applications:

  • Peptide identification in MS-based proteomics: Prioritizes peptides for improved identification in mass spectrometry experiments.
  • Database search space reduction: Reduces the search space during database searches by focusing on likely proteotypic peptides.
  • Large-scale proteomic studies: Supports achieving higher proteome coverage with reduced computational resources in large-scale proteomic analyses.

Methodology:

An SVM model is trained on accurate mass and elution time (AMT) databases derived from tandem mass spectrometry (MS/MS) studies using a descriptor space of 35 amino-acid properties, with demonstrated reduction to 13 variables for prediction of proteotypic peptides.

Topics

Details

Tool Type:
desktop application
Operating Systems:
Windows, Mac
Programming Languages:
Java
Added:
8/3/2017
Last Updated:
11/25/2024

Operations

Publications

Webb-Robertson BM, Cannon WR, Oehmen CS, Shah AR, Gurumoorthi V, Lipton MS, Waters KM. A support vector machine model for the prediction of proteotypic peptides for accurate mass and time proteomics. Bioinformatics. 2010;26(13):1677-1683. doi:10.1093/bioinformatics/btq251. PMID:20568665.

Documentation

Links