PIR

PIR provides curated protein sequence databases and integrated classification resources to support protein annotation, sensitive identification, and comparative analysis in genomics, proteomics, and systems biology.


Key Features:

  • Protein Sequence Database (PSD): An extensively annotated repository containing over 283,000 protein sequences that span the taxonomic spectrum.
  • Family classification methodologies: Employs family classification methods for sensitive identification of proteins and consistent annotation, enabling detection and correction of annotation errors.
  • Superfamily curation and signature domain architectures: Defines signature domain architectures to categorize superfamily membership and enhance automated protein classification accuracy.
  • Bibliography system: Provides literature searching, mapping, and user submission capabilities to augment experimental annotations.
  • Retrospective citation attribution: Performs retrospective attribution of citations for experimental features to enrich annotation provenance.
  • NREF (Non-Redundant Reference Database): Aggregates sequences from PIR-PSD, SWISS-PROT, TrEMBL, RefSeq, GenPept, and the Protein Data Bank (PDB) into a non-redundant reference set exceeding 1 million entries.
  • iProClass integrated database: Integrates protein family, function, and structure information to support systems biology analyses.

Scientific Applications:

  • Genomics and proteomics research: Supplies annotated sequences and classification frameworks to support genome and proteome analyses.
  • Automated protein classification: Enables automated classification using family methodologies and signature domain architectures to improve accuracy and reliability.
  • Annotation curation and enrichment: Facilitates detection and correction of annotation errors and enriches experimental annotations via literature mapping and citation attribution.
  • Systems biology: Provides integrated family, function, and structure data via iProClass for systems-level investigations.

Methodology:

Uses family classification methodologies and superfamily curation to define signature domain architectures; implements a bibliography system for literature searching, mapping, and user submissions; applies retrospective citation attribution to experimental features; and aggregates sequences from PIR-PSD, SWISS-PROT, TrEMBL, RefSeq, GenPept, and PDB to form NREF.

Topics

Collections

Details

Tool Type:
web application
Operating Systems:
Linux, Windows, Mac
Added:
10/7/2015
Last Updated:
11/24/2024

Operations

Publications

Wu CH. The Protein Information Resource. Nucleic Acids Research. 2003;31(1):345-347. doi:10.1093/nar/gkg040. PMID:12520019. PMCID:PMC165487.

Documentation