MzVar

MzVar generates customized variant protein and peptide FASTA databases from VCF variant calls and transcript sequences (e.g., Ensembl, UCSC Table Browser) to enable database searching of mass spectrometry (MS/MS) data.


Key Features:

  • Customized Database Generation: Integrates somatic and genomic variants from VCF files into transcript sequences to compile variant protein and peptide FASTA databases for MS/MS searches.
  • Compatibility with Transcript Sources: Supports transcript sequences from Ensembl and the UCSC Table Browser.
  • Input Formats: Accepts Variant Call Format (VCF) files for variants and FASTA files for transcript sequences.
  • Proteomics Workflow Integration: Generates databases intended for tandem mass spectrometry search workflows (reported usage with X!Tandem) to detect peptides, post-translational modifications, and single amino acid variants.
  • Downstream Analysis Support: Enables downstream filtering at 1% protein-level false discovery rate and spectral counting quantification for evaluation of variant effects on protein expression and stability.

Scientific Applications:

  • Reanalysis of Public Proteomic Datasets: Used to reanalyse 41 publicly available HeLa mass spectrometry datasets.
  • Cell-line Specific Variant Database Compilation: Applied to compile HeLa-specific variant protein databases using somatic and genomic variants from the COSMIC cell line project.
  • Identification of Post-translational Modifications: Facilitates discovery of phosphorylation and N-terminal acetylation sites in MS/MS data searches.
  • Detection of Variant Peptides and Quantitative Comparison: Enables identification of peptides with single amino acid changes and comparison of variant versus wild-type peptide abundance in heterozygous cases.

Methodology:

Uses VCF files and transcript FASTA sequences as inputs; generates variant protein/peptide FASTA databases; searches tandem mass spectra with X!Tandem as reported; filters results at 1% protein-level FDR and applies spectral counting for quantification.

Topics

Collections

Details

Maturity:
Mature
Cost:
Free of charge
Tool Type:
command-line tool
Operating Systems:
Linux, Windows, Mac
Programming Languages:
Java
Added:
9/11/2018
Last Updated:
11/25/2024

Operations

Data Inputs & Outputs

Genetic variation analysis

Inputs

Outputs

Publications

Robin T, Bairoch A, Müller M, Lisacek F, Lane L. Large-Scale Reanalysis of Publicly Available HeLa Cell Proteomics Data in the Context of the Human Proteome Project. Journal of Proteome Research. 2018;17(12):4160-4170. doi:10.1021/acs.jproteome.8b00392. PMID:30175587.