SMAP

SMAP verifies sample identity in proteogenomic studies by detecting variant peptides from multiplexed isobaric labeling-based quantitative mass spectrometry and aligning proteomic-derived allelic information to genomic samples to detect and correct sample mix-ups.


Key Features:

  • Sample identity validation: Leverages quantitative mass spectrometry (MS) data to verify and correct sample identity in large-scale proteogenomic datasets.
  • Proteogenomics variant peptide detection: Detects variant peptides from multiplexed isobaric labeling-based quantitative proteomics data and infers allelic information per sample from peptide expression levels.
  • Discriminant scores for alignment: Computes concordance and specificity discriminant scores to align MS-based proteomic samples with corresponding genomic samples.
  • Theoretical validation with simulations: Uses theoretical analysis with simulation data to demonstrate unique matching capability when at least 20% of individual sample genotypes are available.
  • Empirical cross-validation: Empirically validated in the PsychENCODE BrainGVEX cohort (identified and corrected 54/288 mismatched samples) and cross-validated using ribosome profiling and assay for transposase-accessible chromatin sequencing data from the same samples.

Scientific Applications:

  • Data integrity in proteogenomics: Ensures accurate sample labeling and reduces errors from sample mix-ups in large-scale proteogenomic studies.
  • Integrated genomics–proteomics analyses: Improves validity of studies combining genomic genotypes and quantitative MS-based proteomics by providing sample-level allelic matching.
  • Cohort-level mismatch detection and correction: Detects and corrects sample mismatches in cohort projects exemplified by the PsychENCODE BrainGVEX application.

Methodology:

Detects variant peptides via a proteogenomics approach from multiplexed isobaric labeling quantitative MS, infers allelic information from peptide expression levels, computes concordance and specificity scores to align proteomic and genomic samples, and assesses performance using simulation-based theoretical analysis and empirical validation on the PsychENCODE BrainGVEX dataset with cross-validation against ribosome profiling and assay for transposase-accessible chromatin sequencing data.

Topics

Details

License:
Not licensed
Cost:
Free of charge
Tool Type:
library
Operating Systems:
Linux
Programming Languages:
Perl
Added:
1/28/2022
Last Updated:
1/28/2022

Operations

Data Inputs & Outputs

Genotyping

Publications

Li L, Niu M, Erickson A, Luo J, Rowbotham K, Huang H, Li Y, Jiang Y, Liu C, Peng J, Wang X. SMAP: A pipeline for sample matching in proteogenomics. Unknown Journal. 2021. doi:10.1101/2021.09.17.460682.

Documentation

General', 'User manual
https://smap.shinyapps.io/smap/

Links