XMAn

XMAn provides a non-redundant FASTA database of Homo sapiens mutated peptide sequences for identification and analysis of protein sequence alterations using tandem mass spectrometry (MS/MS).


Key Features:

  • Extensive database: XMAn contains over 2.5 million non-redundant mutated peptide entries derived from 17,599 proteins, including 2,377,103 missense and 161,928 nonsense mutation entries.
  • FASTA format: The dataset is distributed in FASTA format for peptide sequences to support integration with proteomics search engines and MS/MS matching.
  • Mutation-aware proteomics integration: The database supports filtering of protein identification hits to retain only sequences that contain the mutated amino acid for more specific mutation detection.
  • Detailed mutation annotations: Each entry includes protein descriptions and mutation annotations in XMAn.v2 headers (for example, GN=CDC42BPB MRCKB_HUMAN Serine/threonine-protein kinase MRCK beta:c_C3625A:p_L1209I:VGIIEGL:Missense).
  • Non-redundant compilation: Mutated peptide sequences are compiled from multiple sources into a non-redundant set to provide unique entries for analysis.

Scientific Applications:

  • Mutation Analysis: Enable identification and study of specific amino acid changes in Homo sapiens proteins using MS/MS-derived peptide evidence.
  • Disease Research: Support investigation of genetic diseases linked to protein sequence alterations by providing annotated mutation entries.
  • Drug Development: Aid identification of mutations relevant to therapeutic targeting and personalized medicine efforts.

Methodology:

The database was constructed by compiling mutated peptide sequences from multiple sources into a non-redundant FASTA dataset and integrating entries for matching against tandem mass spectrometry (MS/MS) experimental data.

Topics

Details

Programming Languages:
Python
Added:
11/14/2019
Last Updated:
11/24/2024

Operations

Publications

Flores MA, Lazar IM. XMAn v2—a database of <i>Homo sapiens</i> mutated peptides. Bioinformatics. 2019;36(4):1311-1313. doi:10.1093/bioinformatics/btz693. PMID:31539018. PMCID:PMC8215914.

PMID: 31539018
PMCID: PMC8215914
Funding: - National Institute of General Medical Sciences: R01-GM121920