MR-MSPOLYGRAPH

MR-MSPOLYGRAPH performs parallelized peptide–spectrum matching using the MapReduce framework to combine protein sequence database and spectral library searches for improved peptide identification from tandem mass spectra.


Key Features:

  • Parallel Processing with Hadoop: Uses the MapReduce paradigm to distribute peptide–spectrum matching tasks across nodes in a Hadoop cluster, enabling processing of tens of thousands of experimental spectra in hours.
  • Hybrid Search Methodology: Integrates database search (consensus model spectra) and spectral library search (detailed intensity information) to enhance sensitivity and specificity in matching tandem mass spectra.
  • Performance Improvements: Demonstrated substantial speedup and efficiency gains on a 400-core Hadoop cluster using environmental microbial community datasets relative to the serial MSPolygraph implementation.
  • Error Rate Characterization: Employs a scoring metric rooted in statistical thermodynamics to provide conservative error-rate estimates when combining database and spectral-library results, with an estimated false discovery rate of 5%.

Scientific Applications:

  • Global proteomics and large-scale peptide identification: Applied to large datasets to increase the number of spectra assignable to peptides by 57-147%, facilitating protein identification in complex samples such as environmental microbial communities.

Methodology:

Implements MapReduce parallelization on Hadoop; combines protein sequence database searches (consensus model spectra) and spectral library matching (intensity information) for tandem mass spectra; uses a scoring metric based on statistical thermodynamics for error estimation.

Topics

Collections

Details

Tool Type:
command-line tool
Operating Systems:
Linux
Programming Languages:
C
Added:
12/18/2017
Last Updated:
11/24/2024

Operations

Publications

Kalyanaraman A, Cannon WR, Latt B, Baxter DJ. MapReduce implementation of a hybrid spectral library-database search method for large-scale peptide identification. Bioinformatics. 2011;27(21):3072-3073. doi:10.1093/bioinformatics/btr523. PMID:21926122. PMCID:PMC3198583.

Cannon WR, Rawlins MM, Baxter DJ, Callister SJ, Lipton MS, Bryant DA. Large Improvements in MS/MS-Based Peptide Identification Rates using a Hybrid Analysis. Journal of Proteome Research. 2011;10(5):2306-2317. doi:10.1021/pr101130b. PMID:21391700.

Documentation

Downloads

Links