PepExplorer

PepExplorer analyzes de novo peptide sequencing data to assemble homologous protein lists and identify proteins from mass spectrometry experiments when sequence databases are incomplete.


Key Features:

  • De Novo Sequencing Integration: Infers peptide sequences directly from experimental mass spectra without requiring pre-existing protein sequence databases.
  • Pattern Recognition and Protein Assembly: Employs pattern recognition and sequence alignment to assemble homologous protein lists for biological interpretation of de novo peptides.
  • Data Compatibility: Accepts outputs from various de novo sequencing tools used in mass spectrometry workflows.
  • False-Discovery Rate Control: Maintains a global false-discovery rate (FDR) while converging on a list of protein identifications.
  • Neural Network Utilization: Uses a radial basis function neural network that scores precursor charge states, de novo sequencing scores, peptide lengths, and alignment scores to select similar protein candidates from a target-decoy database.
  • Alignment Algorithm: Implements a modified Smith-Waterman algorithm for sequence alignment.
  • Validation and Effectiveness: Recovered most ProLuCID identifications from Pyrococcus furiosus mass spectra against the NCBI RefSeq database, including tests with amino-acid-swapped sequence databases at 1% FDR.
  • Application in Proteomic Research: Enabled a comprehensive proteomic assessment of Bothrops jararaca plasma that identified natural inhibitors of snake toxins.
  • Integration with PatternLab for Proteomics: Integrates with PatternLab for Proteomics to provide access to downstream quantitative and differential proteomics analyses.

Scientific Applications:

  • Novel protein identification: Identifying proteins and peptides when traditional peptide-spectrum matching is limited by incomplete or absent database entries.
  • Homologue assembly and interpretation: Assembling homologous protein lists from de novo peptides to support biological interpretation of complex datasets.
  • Method benchmarking and validation: Validating de novo-based identifications against database search results such as ProLuCID using Pyrococcus furiosus spectra and controlled database manipulations.
  • Complex-sample proteomics: Profiling complex biological samples, exemplified by detection of natural toxin inhibitors in Bothrops jararaca plasma.

Methodology:

Performs de novo peptide inference from mass spectra, applies pattern recognition and a modified Smith-Waterman sequence alignment, scores candidates with a radial basis function neural network using precursor charge, de novo scores, peptide length, and alignment scores against a target-decoy database derived from phylogenetically related species, and applies global FDR control; validation used ProLuCID identifications on Pyrococcus furiosus spectra including amino-acid-swapped database tests.

Topics

Details

Tool Type:
command-line tool
Operating Systems:
Windows
Programming Languages:
C#
Added:
12/18/2017
Last Updated:
1/10/2019

Operations

Data Inputs & Outputs

Protein sequence analysis

Other operations do not define inputs or outputs.

Publications

Leprevost FV, Valente RH, Lima DB, Perales J, Melani R, Yates JR, Barbosa VC, Junqueira M, Carvalho PC. PepExplorer: A Similarity-driven Tool for Analyzing de Novo Sequencing Results. Molecular & Cellular Proteomics. 2014;13(9):2480-2489. doi:10.1074/mcp.m113.037002. PMID:24878498. PMCID:PMC4159663.

PMID: 24878498
PMCID: PMC4159663
Funding: - National Institutes of Health: P41 GM103533, R01 MH067880

Documentation

Links