PfamScan API

PfamScan API annotates catalytic and active-site residues in Pfam protein families by using Pfam Hidden Markov Models (HMMs) to search FASTA sequences and transfer experimentally determined catalytic residue annotations across aligned family members.


Key Features:

  • Pfam HMM FASTA search: Uses Pfam library Hidden Markov Models to search FASTA sequences as the basis for annotation transfer.
  • Active Site Annotation: Predicts and annotates catalytic residues across Pfam families to increase coverage of experimentally characterized active sites.
  • Rule-Based Prediction Methodology: Applies a strict set of rules to transfer experimental catalytic-site annotations between aligned family members while minimizing false positives.
  • Extensive Database Creation: Generates a database containing 606,110 predicted active-site residues, with 94% of these not previously documented in UniProtKB.
  • High Sensitivity and Specificity: Methodology was assessed for sensitivity and specificity, estimating approximately 3% false positives among predictions.
  • Comparative Analysis: Compares predictions with UniProtKB, Catalytic Site Atlas, PROSITE, and MEROPS to identify novel predictions and limitations in other resources.
  • Scalability and Flexibility: Applies the data-transfer methodology to any alignment that has associated experimental active-site information.
  • Regular Updates: Recalculates active-site predictions at each Pfam release to maintain currency of the annotation set.

Scientific Applications:

  • Enzyme Function Prediction: Identifies catalytic residues in uncharacterized proteins within enzymatic Pfam families to support inference of enzymatic functions.
  • Protein Engineering and Drug Design: Provides residue-level active-site annotations to inform protein engineering strategies and the design of inhibitors or activators targeting specific catalytic residues.
  • Comparative Genomics: Enables analysis of active-site conservation across species and the study of evolutionary relationships among proteins.

Methodology:

Searches FASTA sequences with Pfam HMMs and transfers experimentally determined catalytic residue annotations across aligned family members using a strict rule set; predictions were compiled into a database, assessed for sensitivity and specificity (≈3% estimated false positives), and are recalculated at each Pfam release.

Topics

Details

Tool Type:
api
Operating Systems:
Linux, Windows, Mac
Added:
8/3/2015
Last Updated:
11/24/2024

Operations

Publications

Mistry J, Bateman A, Finn RD. Predicting active site residue annotations in the Pfam database. BMC Bioinformatics. 2007;8(1). doi:10.1186/1471-2105-8-298. PMID:17688688. PMCID:PMC2025603.

Documentation

Links