SPROF-GO

SPROF-GO predicts protein functions directly from amino acid sequences using pretrained language model embeddings, self-attention pooling, and homology-based label diffusion without relying on structural or network data.


Key Features:

  • Pretrained Language Model Integration: A pretrained language model extracts informative sequence embeddings that capture intricate patterns within protein sequences.
  • Self-Attention Pooling Mechanism: Self-attention pooling focuses on crucial residues within sequences to identify regions most informative for function prediction.
  • Homology-Based Label Diffusion: Incorporates homology information and label diffusion algorithms that consider overlapping communities of proteins to refine function predictions without direct structural or network data.
  • Performance Superiority: Demonstrated improvements of over 14.5%, 27.3%, and 10.1% in the area under the precision-recall curve across three sub-ontology test sets compared to state-of-the-art sequence-based and network-based methods.
  • Generalization Capabilities: Generalizes across non-homologous proteins and across species not encountered during training.
  • Visualization of Attention Mechanisms: Provides attention-based visualizations to indicate sequence domains contributing most to predicted functions.

Scientific Applications:

  • Elucidating Disease Mechanisms: Supports functional characterization of proteins relevant to disease mechanisms by predicting likely functions from sequence alone.
  • Identifying Potential Drug Targets: Aids identification and prioritization of candidate drug targets through predicted protein functions.
  • Functional Annotation Across Diverse Species: Enables annotation of proteins in species or datasets lacking structural or network information by relying solely on sequence data.

Methodology:

Uses a pretrained language model to extract sequence embeddings, applies self-attention pooling to weight informative residues, and refines predictions via homology-based label diffusion across overlapping protein communities.

Topics

Details

License:
MIT
Cost:
Free of charge
Tool Type:
web application
Operating Systems:
Mac, Linux, Windows
Programming Languages:
Python
Added:
11/7/2023
Last Updated:
11/24/2024

Operations

Publications

Yuan Q, Xie J, Xie J, Zhao H, Yang Y. Fast and accurate protein function prediction from sequence through pretrained language model and homology-based label diffusion. Briefings in Bioinformatics. 2023;24(3). doi:10.1093/bib/bbad117. PMID:36964722.

PMID: 36964722
Funding: - Guangzhou S&T Research Plan: 202002020047, 202007030010 - Guangdong Key Field R&D Plan: 2018B0101090060, 2019B020228001 - National Natural Science Foundation of China: 12126610 - National Key R&D Program of China: 2022YFF1203100

Links