GFS

GFS maps peptide mass fingerprints to genomic sequences to enable annotation-independent protein identification.


Key Features:

  • Genome-wide theoretical translation and digest scanning: Scans peptide mass fingerprints against the theoretical translation and proteolytic digest of an entire genome sequence.
  • Annotation-independent mapping: Operates without predefined open reading frames (ORFs) or existing protein annotations.
  • Windowed scoring system: Scores fixed-size windows along the genome using number of matching peptides, missed enzymatic cleavages, presence of in-frame stop codons, peptide adjacency, and duplicate matches.
  • Statistical significance assessment: Assesses significance by comparing window scores to scores from windows matched with randomly produced mass data.
  • Validation: Validated on Saccharomyces cerevisiae mitochondria and Escherichia coli with 86% concordance to peptident and mascot.
  • Annotation error detection: Identifies proteins with incorrect or missing annotations and proteins affected by sequencing-induced framing mistakes.
  • Recoding event detection: Detects proteins affected by recoding events such as frameshifting and stop-codon read-through.
  • Distributed computation: Implemented as a client-server pair enabling distribution and concurrent analysis across cluster nodes.

Scientific Applications:

  • Genome annotation: Locates peptide-supported coding regions independent of existing annotations.
  • Proteomics with incomplete databases: Enables protein identification when protein databases lag behind genome sequencing and annotation.
  • Annotation correction: Detects and supports correction of incorrect or missing protein annotations.
  • Sequencing error analysis: Reveals proteins impacted by sequencing errors that cause framing mistakes.
  • Recoding studies: Identifies candidates for frameshifting and stop-codon read-through investigations.

Methodology:

Scans peptide mass fingerprints against the theoretical translation and proteolytic digest of an entire genome; evaluates fixed-size genomic windows by matching peptide counts, missed enzymatic cleavages, in-frame stop codons, peptide adjacency, and duplicate matches; assesses statistical significance by comparing window scores to scores from windows matched with randomly produced mass data; implemented as a client-server pair for distributed concurrent analysis across cluster nodes.

Topics

Collections

Details

Tool Type:
web application
Added:
4/21/2017
Last Updated:
11/25/2024

Operations

Publications

Wisz MS, Suarez MK, Holmes MR, Giddings MC. GFSWeb:  A Web Tool for Genome-Based Identification of Proteins from Mass Spectrometric Samples. Journal of Proteome Research. 2004;3(6):1292-1295. doi:10.1021/pr049879y. PMID:15595741. PMCID:PMC1351070.

Giddings MC, Shah AA, Gesteland R, Moore B. Genome-based peptide fingerprint scanning. Proceedings of the National Academy of Sciences. 2002;100(1):20-25. doi:10.1073/pnas.0136893100. PMID:12518051. PMCID:PMC140871.