STEPP
STEPP predicts proteotypic peptides for mass spectrometry-based proteomics using a support vector machine to improve peptide identification.
Key Features:
- Quantitative Prediction: Quantitatively predicts proteotypic peptides to prioritize peptides most likely to be detected by mass spectrometry.
- Support Vector Machine model: Employs a support vector machine (SVM) that uses a descriptor space derived from 35 amino-acid properties including content, charge, hydrophilicity, and polarity.
- Validation on AMT databases: Validated on three independently derived accurate mass and elution time (AMT) databases from Shewanella oneidensis, Salmonella typhimurium, and Yersinia pestis, achieving an average accuracy of approximately 0.83 with standard deviation < 0.038.
- Reduced descriptor set: Maintains high prediction accuracy when the descriptor space is reduced to a small set of 13 variables.
- Cross-species validation: Demonstrates robustness across different species used in validation, supporting applicability in diverse proteomic datasets.
Scientific Applications:
- Peptide identification in MS-based proteomics: Prioritizes peptides for improved identification in mass spectrometry experiments.
- Database search space reduction: Reduces the search space during database searches by focusing on likely proteotypic peptides.
- Large-scale proteomic studies: Supports achieving higher proteome coverage with reduced computational resources in large-scale proteomic analyses.
Methodology:
An SVM model is trained on accurate mass and elution time (AMT) databases derived from tandem mass spectrometry (MS/MS) studies using a descriptor space of 35 amino-acid properties, with demonstrated reduction to 13 variables for prediction of proteotypic peptides.
Topics
Details
- Tool Type:
- desktop application
- Operating Systems:
- Windows, Mac
- Programming Languages:
- Java
- Added:
- 8/3/2017
- Last Updated:
- 11/25/2024
Operations
Publications
Webb-Robertson BM, Cannon WR, Oehmen CS, Shah AR, Gurumoorthi V, Lipton MS, Waters KM. A support vector machine model for the prediction of proteotypic peptides for accurate mass and time proteomics. Bioinformatics. 2010;26(13):1677-1683. doi:10.1093/bioinformatics/btq251. PMID:20568665.
PMID: 20568665