DisoFLAG

DisoFLAG predicts protein intrinsic disorder and associated functions from amino acid sequences to characterize intrinsically disordered proteins and regions (IDPs/IDRs).


Key Features:

  • Graph-Based Interaction Protein Language Model (GiPLM): DisoFLAG integrates semantic information from pre-trained protein language models into graph-based interaction units to enhance correlation of semantic representations across multiple disordered functions.
  • Comprehensive Function Prediction: DisoFLAG predicts six disordered functions—Protein-binding, DNA-binding, RNA-binding, Ion-binding, Lipid-binding, and Flexible linker—using sequence information alone.
  • Sequence-only input: DisoFLAG requires only amino acid sequences as input to generate disorder and function predictions.
  • High Predictive Accuracy: DisoFLAG's performance was evaluated in Critical Assessment of protein Intrinsic Disorder (CAID) experiments and demonstrated accurate and comprehensive predictions.

Scientific Applications:

  • IDP/IDR functional annotation: Predicting intrinsic disorder and the six associated functions to aid characterization of intrinsically disordered proteins and regions in molecular studies.
  • Protein interaction and mechanism inference: Informing studies of protein interactions and cellular mechanisms by identifying disorder-mediated binding and flexible linker regions.

Methodology:

DisoFLAG takes amino acid sequences as input and leverages the GiPLM framework, integrating semantic representations from pre-trained protein language models into graph-based interaction units to predict intrinsic disorder and associated functions.

Topics

Details

License:
MIT
Cost:
Free of charge
Tool Type:
workflow
Operating Systems:
Mac, Linux, Windows
Programming Languages:
Python
Added:
5/18/2024
Last Updated:
11/24/2024

Operations

Data Inputs & Outputs

DNA-binding protein prediction

Publications

Pang Y, Liu B. DisoFLAG: accurate prediction of protein intrinsic disorder and its functions using graph-based interaction protein language model. BMC Biology. 2024;22(1). doi:10.1186/s12915-023-01803-y. PMID:38166858. PMCID:PMC10762911.

PMID: 38166858
Funding: - National Natural Science Foundation of China: 62250028, 62271049, 62325202, U22A2039