SpeCollate

SpeCollate learns cross-modal similarity between experimental mass spectra and peptide sequences using a deep cross-modal similarity network to improve peptide identification accuracy in MS-based proteomics.


Key Features:

  • Deep learning approach: Employs a deep cross-modal similarity network trained on labeled MS data to learn the similarity function between experimental spectra and peptides.
  • Shared Euclidean subspace transformation: Transforms experimental spectra and peptide sequences into fixed-size embeddings in a shared Euclidean subspace for direct similarity comparison.
  • Custom SNAP-loss function: Optimizes network training with a custom SNAP-loss function tailored to discriminate subtle similarities between spectra and peptides.
  • Online hardest negative mining: Incorporates online hardest negative mining to select challenging negative examples during training.
  • Extensive training dataset: Trained on 4.8 million sextuplets derived from the NIST and MassIVE peptide libraries.

Scientific Applications:

  • Peptide identification: Improves peptide-spectrum match (PSM) accuracy and peptide identification in MS-based proteomics.
  • Enhanced recovery at low FDR: Recovers more PSMs and unique peptides than Crux and MSFragger, including under stringent false discovery rate (FDR) thresholds (<1%).
  • Novel peptide discovery: Identifies peptides not reported by existing database search methods, enabling additional peptide identifications.

Methodology:

The method trains a deep cross-modal network on 4.8 million sextuplets from NIST and MassIVE, learning a similarity function from labeled MS data by transforming spectra and peptides into a shared Euclidean subspace and optimizing with a custom SNAP-loss and online hardest negative mining.

Topics

Details

License:
Not licensed
Cost:
Free of charge
Tool Type:
command-line tool
Operating Systems:
Linux
Added:
3/13/2022
Last Updated:
3/13/2022

Operations

Publications

Tariq MU, Saeed F. SpeCollate: Deep cross-modal similarity network for mass spectrometry data based peptide deductions. PLOS ONE. 2021;16(10):e0259349. doi:10.1371/journal.pone.0259349. PMID:34714871. PMCID:PMC8555789.

PMID: 34714871
PMCID: PMC8555789
Funding: - Division of Advanced Cyberinfrastructure: OAC 1925960 - National Institute of General Medical Sciences: R01GM134384