Entrez Protein Clusters

Entrez Protein Clusters organizes protein sequences into clusters by sequence similarity to support analysis of protein function and evolutionary relationships across prokaryotes, bacteriophages, mitochondria, and chloroplasts.


Key Features:

  • Data Aggregation and Organization: The Protein Clusters Database (ProtClustDB) groups proteins into clusters based on sequence similarity to enable study of protein functions and evolutionary relationships.
  • Comprehensive Dataset: As of May 2008, ProtClustDB contains 285,386 clusters derived from over 1.7 million proteins encoded by 3,806 nucleotide sequences from the RefSeq collection, including complete chromosomes and plasmids for prokaryotes, bacteriophages, mitochondria, and chloroplasts.
  • Curated Functional Annotation: A subset of 7,180 clusters comprises 376,513 proteins with curated gene and protein functional annotations.
  • Integration with External Resources: Clusters are supplemented with PubMed identifiers and external cross-references for linkage to literature and other databases.
  • Advanced Analysis Capabilities: Provides generation of multiple sequence alignments, construction of phylogenetic trees, and examination of genomic neighborhoods for in-depth comparative and evolutionary analyses.

Scientific Applications:

  • Evolutionary Biology: Enables inference of phylogenetic relationships and comparative evolutionary studies across prokaryotes, bacteriophages, mitochondria, and chloroplasts.
  • Functional Annotation and Comparative Genomics: Supports assignment and comparison of gene and protein functions using clustered sequences and curated annotations.
  • Microbiology and Microbial Genomics: Facilitates analysis of microbial protein function, genomic context, and evolutionary dynamics in bacteria and phages.

Methodology:

Protein sequences are grouped into clusters by sequence similarity, and the database structure supports updates and maintenance.

Topics

Collections

Details

Tool Type:
web application
Operating Systems:
Linux, Windows, Mac
Added:
9/12/2015
Last Updated:
11/25/2024

Operations

Data Inputs & Outputs

Query and retrieval

Publications

Klimke W, Agarwala R, Badretdin A, Chetvernin S, Ciufo S, Fedorov B, Kiryutin B, O’Neill K, Resch W, Resenchuk S, Schafer S, Tolstoy I, Tatusova T. The National Center for Biotechnology Information's Protein Clusters Database. Nucleic Acids Research. 2008;37(suppl_1):D216-D223. doi:10.1093/nar/gkn734. PMID:18940865. PMCID:PMC2686591.

Documentation