SAND

SAND performs automated decomposition and quantification of NMR spectra to enable accurate metabolite quantification in untargeted NMR metabolomics.


Key Features:

  • Automated Time-Domain Modeling: Performs automatic time-domain modeling to quantify entire NMR spectra without manual intervention.
  • Hybrid Optimization Techniques: Integrates Markov chain Monte Carlo (MCMC) with subsampling in both time and frequency domains to improve robustness and accuracy of spectral decomposition.
  • Overfitting Mitigation: Randomly divides time-domain data into training and validation sets to reduce overfitting and improve generalization.
  • High Accuracy and Correlation: Achieves a correlation coefficient of approximately 0.9 with ground truth across highly overlapped simulated datasets, two-compound mixtures, and urine samples spiked with a four-compound mixture.
  • Automated Annotation via Correlation Networks: Constructs correlation networks from decomposed spectral peaks and recovers approximately 74% of peaks per compound within single clusters.
  • Scalability and Flexibility: Relies on time-domain subsampling using random subsets of time-domain points, enabling extensions to higher-dimensional data and nonuniformly sampled datasets.

Scientific Applications:

  • Untargeted metabolomics: Supports quantitative profiling of complex biological sample sets through automated spectral decomposition and quantification.
  • Metabolite identification and characterization: Facilitates recovery and annotation of compound-specific peaks via correlation networks from decomposed spectra.
  • Biomarker discovery and pathway analysis: Provides quantitative spectral measures suitable for metabolic pathway analysis and biomarker discovery in biological samples such as urine.

Methodology:

Automatic time-domain modeling and automated quantification of entire NMR spectra; hybrid optimization combining Markov chain Monte Carlo (MCMC) with subsampling in time and frequency domains; random time-domain subsampling using random subsets of points; random division of time-domain data into training and validation sets to mitigate overfitting; construction of correlation networks from decomposed spectral peaks for automated annotation.

Topics

Details

License:
MIT
Cost:
Free of charge
Tool Type:
workflow
Programming Languages:
R, MATLAB
Added:
5/24/2024
Last Updated:
11/24/2024

Operations

Publications

Wu Y, Sanati O, Uchimiya M, Krishnamurthy K, Wedell J, Hoch JC, Edison AS, Delaglio F. SAND: Automated Time-Domain Modeling of NMR Spectra Applied to Metabolite Quantification. Analytical Chemistry. 2024;96(5):1843-1851. doi:10.1021/acs.analchem.3c03078. PMID:38273718. PMCID:PMC10896553.

PMID: 38273718
Funding: - Division of Biological Infrastructure: 1946970