SAND
SAND performs automated decomposition and quantification of NMR spectra to enable accurate metabolite quantification in untargeted NMR metabolomics.
Key Features:
- Automated Time-Domain Modeling: Performs automatic time-domain modeling to quantify entire NMR spectra without manual intervention.
- Hybrid Optimization Techniques: Integrates Markov chain Monte Carlo (MCMC) with subsampling in both time and frequency domains to improve robustness and accuracy of spectral decomposition.
- Overfitting Mitigation: Randomly divides time-domain data into training and validation sets to reduce overfitting and improve generalization.
- High Accuracy and Correlation: Achieves a correlation coefficient of approximately 0.9 with ground truth across highly overlapped simulated datasets, two-compound mixtures, and urine samples spiked with a four-compound mixture.
- Automated Annotation via Correlation Networks: Constructs correlation networks from decomposed spectral peaks and recovers approximately 74% of peaks per compound within single clusters.
- Scalability and Flexibility: Relies on time-domain subsampling using random subsets of time-domain points, enabling extensions to higher-dimensional data and nonuniformly sampled datasets.
Scientific Applications:
- Untargeted metabolomics: Supports quantitative profiling of complex biological sample sets through automated spectral decomposition and quantification.
- Metabolite identification and characterization: Facilitates recovery and annotation of compound-specific peaks via correlation networks from decomposed spectra.
- Biomarker discovery and pathway analysis: Provides quantitative spectral measures suitable for metabolic pathway analysis and biomarker discovery in biological samples such as urine.
Methodology:
Automatic time-domain modeling and automated quantification of entire NMR spectra; hybrid optimization combining Markov chain Monte Carlo (MCMC) with subsampling in time and frequency domains; random time-domain subsampling using random subsets of points; random division of time-domain data into training and validation sets to mitigate overfitting; construction of correlation networks from decomposed spectral peaks for automated annotation.
Topics
Details
- License:
- MIT
- Cost:
- Free of charge
- Tool Type:
- workflow
- Programming Languages:
- R, MATLAB
- Added:
- 5/24/2024
- Last Updated:
- 11/24/2024
Operations
Publications
Wu Y, Sanati O, Uchimiya M, Krishnamurthy K, Wedell J, Hoch JC, Edison AS, Delaglio F. SAND: Automated Time-Domain Modeling of NMR Spectra Applied to Metabolite Quantification. Analytical Chemistry. 2024;96(5):1843-1851. doi:10.1021/acs.analchem.3c03078. PMID:38273718. PMCID:PMC10896553.