ProtoNet

ProtoNet classifies protein sequences into a hierarchical tree to represent protein families across the protein sequence space.


Key Features:

  • Unsupervised bottom-up clustering: Uses an unsupervised bottom-up clustering algorithm to organize large protein datasets into a hierarchical tree of clusters.
  • All-against-all sequence comparison: Constructs the hierarchical tree based on an all-against-all comparison using 2.5 million representative sequences from UniRef50.
  • Database coverage: Includes over 9 million proteins with 5.5% from UniProtKB/SwissProt and the remainder from UniProtKB/TrEMBL.
  • Annotation-based quality filtering: Applies rigorous annotation-based tests to identify a refined set of 162,088 high-quality clusters.
  • ProtoName assignment: Assigns each cluster a unique ProtoName derived from Gene Ontology (GO) terms, UniProt/Swiss-Prot keywords, and InterPro entries.
  • Multi-resolution analysis modes: Provides a default overview of the entire hierarchy and an advanced mode to explore the family tree at varying levels of resolution.
  • Version comparison capability: Enables systematic comparison of different ProtoNet versions to observe changes as sequence space knowledge increases.

Scientific Applications:

  • Protein family classification: Classifies proteins into hierarchical families to facilitate the study of protein relationships.
  • Functional and structural annotation: Supports annotation of clusters via ProtoName assignments using GO, UniProt/Swiss-Prot keywords, and InterPro to identify functional and structural characteristics.
  • Large-scale proteome analysis: Enables analysis of proteome-scale data drawn from UniProtKB/SwissProt and UniProtKB/TrEMBL.
  • Comparative clustering analysis: Facilitates comparison of clustering outcomes across ProtoNet versions to assess how family definitions change with expanded sequence data.

Methodology:

Hierarchical tree construction via an all-against-all comparison of 2.5 million UniRef50 representative sequences, unsupervised bottom-up clustering to form clusters, annotation-based tests to filter for 162,088 high-quality clusters, and assignment of ProtoNames from GO terms, UniProt/Swiss-Prot keywords, and InterPro entries.

Topics

Collections

Details

Tool Type:
web application
Operating Systems:
Linux, Windows, Mac
Added:
10/7/2015
Last Updated:
11/25/2024

Operations

Publications

Rappoport N, Karsenty S, Stern A, Linial N, Linial M. ProtoNet 6.0: organizing 10 million protein sequences in a compact hierarchical family tree. Nucleic Acids Research. 2011;40(D1):D313-D320. doi:10.1093/nar/gkr1027. PMID:22121228. PMCID:PMC3245180.

Documentation