MR-MSPOLYGRAPH
MR-MSPOLYGRAPH performs parallelized peptide–spectrum matching using the MapReduce framework to combine protein sequence database and spectral library searches for improved peptide identification from tandem mass spectra.
Key Features:
- Parallel Processing with Hadoop: Uses the MapReduce paradigm to distribute peptide–spectrum matching tasks across nodes in a Hadoop cluster, enabling processing of tens of thousands of experimental spectra in hours.
- Hybrid Search Methodology: Integrates database search (consensus model spectra) and spectral library search (detailed intensity information) to enhance sensitivity and specificity in matching tandem mass spectra.
- Performance Improvements: Demonstrated substantial speedup and efficiency gains on a 400-core Hadoop cluster using environmental microbial community datasets relative to the serial MSPolygraph implementation.
- Error Rate Characterization: Employs a scoring metric rooted in statistical thermodynamics to provide conservative error-rate estimates when combining database and spectral-library results, with an estimated false discovery rate of 5%.
Scientific Applications:
- Global proteomics and large-scale peptide identification: Applied to large datasets to increase the number of spectra assignable to peptides by 57-147%, facilitating protein identification in complex samples such as environmental microbial communities.
Methodology:
Implements MapReduce parallelization on Hadoop; combines protein sequence database searches (consensus model spectra) and spectral library matching (intensity information) for tandem mass spectra; uses a scoring metric based on statistical thermodynamics for error estimation.
Topics
Collections
Details
- Tool Type:
- command-line tool
- Operating Systems:
- Linux
- Programming Languages:
- C
- Added:
- 12/18/2017
- Last Updated:
- 11/24/2024
Operations
Publications
Kalyanaraman A, Cannon WR, Latt B, Baxter DJ. MapReduce implementation of a hybrid spectral library-database search method for large-scale peptide identification. Bioinformatics. 2011;27(21):3072-3073. doi:10.1093/bioinformatics/btr523. PMID:21926122. PMCID:PMC3198583.
Cannon WR, Rawlins MM, Baxter DJ, Callister SJ, Lipton MS, Bryant DA. Large Improvements in MS/MS-Based Peptide Identification Rates using a Hybrid Analysis. Journal of Proteome Research. 2011;10(5):2306-2317. doi:10.1021/pr101130b. PMID:21391700.
Documentation
Downloads
- Source codehttp://compbio.eecs.wsu.edu/