NOmESS

NOmESS constructs non-redundant reference protein sequence sets for mass spectrometry-based proteomics by homology-driven assembly of translated expressed sequence tags (ESTs) and RNA deep-sequencing data aligned to proteins from a closely related, fully sequenced organism.


Key Features:

  • Homology-Driven Assembly: Leverages translated DNA sequences (ESTs and RNA deep-sequencing) aligned to a closely related organism's proteins to infer protein-coding regions.
  • Non-Redundant Reference Set: Clusters similar sequences and selects representative sequences to maximize amino acid sequence information and produce a non-redundant dataset.
  • Sequence Alignment and Clustering: Uses BLASTp for sequence alignment and cd-hit for sequence clustering.
  • Implementation: Implemented in C#.

Scientific Applications:

  • MS-based Protein Identification: Provides reference databases to improve protein identification and characterization in mass spectrometry-based proteomics.
  • Poorly Characterized Organisms: Enables proteomic analyses for organisms with limited genomic resources by generating protein sequence resources from ESTs and RNA deep-sequencing.
  • Xenopus laevis Reference Assembly: Has been applied to assemble a reference database for Xenopus laevis.

Methodology:

Translated ESTs and RNA deep-sequencing data are aligned to proteins from a closely related fully sequenced organism using BLASTp; similar sequences are clustered with cd-hit and representative sequences selected to form the non-redundant protein set; the software is implemented in C#.

Topics

Collections

Details

Tool Type:
command-line tool
Operating Systems:
Windows
Programming Languages:
C#
Added:
8/3/2017
Last Updated:
11/25/2024

Operations

Publications

Temu T, Mann M, Räschle M, Cox J. Homology-driven assembly of NOn-redundant protEin sequence sets (NOmESS) for mass spectrometry. Bioinformatics. 2015;32(9):1417-1419. doi:10.1093/bioinformatics/btv756. PMID:26743511. PMCID:PMC4848398.

Documentation

Links