SF-Matching

SF-Matching predicts metabolite structures from tandem mass spectrometry (MS/MS) spectra by matching substructure-associated fragmentation patterns using machine-learning random forest models.


Key Features:

  • SubFragment-Matching Hypothesis: Assumes molecules sharing substructures produce analogous fragmentation patterns that can be exploited for identification from MS/MS spectra.
  • Integration of Fragmentation Patterns and Substructures: Analyzes fragmentation patterns linked to shared molecular substructures to estimate the likelihood that a candidate structure generates observed fragments.
  • Machine Learning Model - Random Forests: Employs random forest models to score candidate molecules based on their predicted fragmentation behavior.
  • Pre-computation of Scores: Pre-computes scores across common biological molecular structure databases to enable rapid comparison of experimental spectra to candidate structures.
  • Benchmarking and Performance: Benchmarked on benchmark datasets showing performance comparable to CSI: FingerID and complementary accuracy gains when combined with CSI: FingerID.
  • Scalability and Data-Driven Improvement: Rarefaction analysis of the training dataset indicates performance improves as additional experimental MS/MS data are incorporated.

Scientific Applications:

  • Untargeted metabolomics: Supports identification workflows for untargeted MS/MS data by predicting compound identities from complex spectra.
  • Metabolic pathway elucidation: Aids elucidation of metabolic pathways via structural annotation of detected metabolites.
  • Biomarker discovery: Facilitates discovery of novel biomarkers through structure-level annotation of MS/MS features.
  • Biochemical process investigation: Enables studies of biochemical processes by providing compound-level identifications from MS/MS data.

Methodology:

Matches substructure-associated fragmentation patterns, analyzes fragmentation patterns and substructures, uses random forest models to assess and score candidate molecules, pre-computes scores across common biological molecular structure databases, and performs rarefaction analysis on the training dataset.

Topics

Details

Added:
1/9/2020
Last Updated:
1/16/2021

Operations

Publications

Li Y, Kuhn M, Gavin A, Bork P. Identification of metabolites from tandem mass spectra with a machine learning approach utilizing structural features. Bioinformatics. 2019;36(4):1213-1218. doi:10.1093/bioinformatics/btz736. PMID:31605112. PMCID:PMC7703789.

PMID: 31605112
Funding: - EMBL and the MicrobioS: ERC-AdG-669830

Links