MzVar
MzVar generates customized variant protein and peptide FASTA databases from VCF variant calls and transcript sequences (e.g., Ensembl, UCSC Table Browser) to enable database searching of mass spectrometry (MS/MS) data.
Key Features:
- Customized Database Generation: Integrates somatic and genomic variants from VCF files into transcript sequences to compile variant protein and peptide FASTA databases for MS/MS searches.
- Compatibility with Transcript Sources: Supports transcript sequences from Ensembl and the UCSC Table Browser.
- Input Formats: Accepts Variant Call Format (VCF) files for variants and FASTA files for transcript sequences.
- Proteomics Workflow Integration: Generates databases intended for tandem mass spectrometry search workflows (reported usage with X!Tandem) to detect peptides, post-translational modifications, and single amino acid variants.
- Downstream Analysis Support: Enables downstream filtering at 1% protein-level false discovery rate and spectral counting quantification for evaluation of variant effects on protein expression and stability.
Scientific Applications:
- Reanalysis of Public Proteomic Datasets: Used to reanalyse 41 publicly available HeLa mass spectrometry datasets.
- Cell-line Specific Variant Database Compilation: Applied to compile HeLa-specific variant protein databases using somatic and genomic variants from the COSMIC cell line project.
- Identification of Post-translational Modifications: Facilitates discovery of phosphorylation and N-terminal acetylation sites in MS/MS data searches.
- Detection of Variant Peptides and Quantitative Comparison: Enables identification of peptides with single amino acid changes and comparison of variant versus wild-type peptide abundance in heterozygous cases.
Methodology:
Uses VCF files and transcript FASTA sequences as inputs; generates variant protein/peptide FASTA databases; searches tandem mass spectra with X!Tandem as reported; filters results at 1% protein-level FDR and applies spectral counting for quantification.
Topics
Collections
Details
- Maturity:
- Mature
- Cost:
- Free of charge
- Tool Type:
- command-line tool
- Operating Systems:
- Linux, Windows, Mac
- Programming Languages:
- Java
- Added:
- 9/11/2018
- Last Updated:
- 11/25/2024
Operations
Data Inputs & Outputs
Publications
Robin T, Bairoch A, Müller M, Lisacek F, Lane L. Large-Scale Reanalysis of Publicly Available HeLa Cell Proteomics Data in the Context of the Human Proteome Project. Journal of Proteome Research. 2018;17(12):4160-4170. doi:10.1021/acs.jproteome.8b00392. PMID:30175587.