ATGO

ATGO predicts protein Gene Ontology (GO) attributes from protein sequences using self-attention transformer-based pre-trained language models combined with triplet neural networks.


Key Features:

  • Pre-trained transformer language models: Employs self-attention transformer-based pre-trained language models on protein sequences to extract discriminative functional patterns from feature embeddings.
  • Triplet neural network: Uses a triplet neural network architecture to align functional similarity with embedding feature similarity for refined GO prediction.
  • GO domain coverage: Predicts GO attributes across molecular function, biological process, and cellular component.
  • Benchmarking and validation: Evaluated on 1,068 non-redundant benchmarking proteins and 3,328 targets from the third Critical Assessment of Protein Function Annotation (CAFA) challenge.
  • Integration with homology and network scores: Combines model outputs with network scores and complementary homology-based inferences to improve prediction accuracy.
  • Improved prediction accuracy: Demonstrates superior performance over existing state-of-the-art approaches in GO prediction as reported in validation studies.

Scientific Applications:

  • Protein function annotation: Enables high-accuracy assignment of GO terms to proteins based on sequence-derived embeddings.
  • Proteome-scale annotation: Supports large-scale automated protein function annotation from sequence data.
  • Drug target characterization: Facilitates elucidation of protein functions relevant to drug design.
  • Functional genomics and proteomics: Improves interpretation of molecular functions, biological processes, and cellular localization in genomics and proteomics analyses.

Methodology:

Pre-trained self-attention transformer language models generate protein sequence feature embeddings; a triplet neural network is trained on these embeddings to align functional similarity with feature similarity; model outputs can be combined with network scores and complementary homology-based inference; methods were evaluated on 1,068 non-redundant benchmarking proteins and 3,328 targets from the third Critical Assessment of Protein Function Annotation (CAFA) challenge.

Topics

Details

Cost:
Free of charge
Tool Type:
web application
Operating Systems:
Mac, Linux, Windows
Added:
2/28/2023
Last Updated:
11/24/2024

Operations

Data Inputs & Outputs

Publications

Zhu Y, Zhang C, Yu D, Zhang Y. Integrating unsupervised language model with triplet neural networks for protein gene ontology prediction. PLOS Computational Biology. 2022;18(12):e1010793. doi:10.1371/journal.pcbi.1010793. PMID:36548439. PMCID:PMC9822105.

PMID: 36548439
PMCID: PMC9822105
Funding: - China Scholarship Council: 201906840041 - National Natural Science Foundation of China: 61772273, 62072243 - Natural Science Foundation of Jiangsu: BK20201304 - Foundation of National Defense Key Laboratory of Science and Technology: JZX7Y202001SY000901 - National Institute of General Medical Sciences: GM136422, S10OD026825 - National Institute of Allergy and Infectious Diseases: AI134678 - National Science Foundation: DBI2030790, IIS1901191, MTM2025426