PIR
PIR provides curated protein sequence databases and integrated classification resources to support protein annotation, sensitive identification, and comparative analysis in genomics, proteomics, and systems biology.
Key Features:
- Protein Sequence Database (PSD): An extensively annotated repository containing over 283,000 protein sequences that span the taxonomic spectrum.
- Family classification methodologies: Employs family classification methods for sensitive identification of proteins and consistent annotation, enabling detection and correction of annotation errors.
- Superfamily curation and signature domain architectures: Defines signature domain architectures to categorize superfamily membership and enhance automated protein classification accuracy.
- Bibliography system: Provides literature searching, mapping, and user submission capabilities to augment experimental annotations.
- Retrospective citation attribution: Performs retrospective attribution of citations for experimental features to enrich annotation provenance.
- NREF (Non-Redundant Reference Database): Aggregates sequences from PIR-PSD, SWISS-PROT, TrEMBL, RefSeq, GenPept, and the Protein Data Bank (PDB) into a non-redundant reference set exceeding 1 million entries.
- iProClass integrated database: Integrates protein family, function, and structure information to support systems biology analyses.
Scientific Applications:
- Genomics and proteomics research: Supplies annotated sequences and classification frameworks to support genome and proteome analyses.
- Automated protein classification: Enables automated classification using family methodologies and signature domain architectures to improve accuracy and reliability.
- Annotation curation and enrichment: Facilitates detection and correction of annotation errors and enriches experimental annotations via literature mapping and citation attribution.
- Systems biology: Provides integrated family, function, and structure data via iProClass for systems-level investigations.
Methodology:
Uses family classification methodologies and superfamily curation to define signature domain architectures; implements a bibliography system for literature searching, mapping, and user submissions; applies retrospective citation attribution to experimental features; and aggregates sequences from PIR-PSD, SWISS-PROT, TrEMBL, RefSeq, GenPept, and PDB to form NREF.
Topics
Collections
Details
- Tool Type:
- web application
- Operating Systems:
- Linux, Windows, Mac
- Added:
- 10/7/2015
- Last Updated:
- 11/24/2024
Operations
Publications
Wu CH. The Protein Information Resource. Nucleic Acids Research. 2003;31(1):345-347. doi:10.1093/nar/gkg040. PMID:12520019. PMCID:PMC165487.