DMFpred

DMFpred predicts molecular functions of intrinsically disordered proteins and regions (IDPs/IDRs) at the residue level to identify functionally annotated residues relevant to drug target discovery and rational drug design.


Key Features:

  • Function scope: Predicts six disordered molecular function types: entropic chain, assembler, scavenger, effector, display site, and chaperone.
  • Residue-level and multi-functional prediction: Provides residue-level predictions and supports identification of functional and multi-functional residues across five evaluation categories.
  • Protein Cubic Language Model (PCLM): Implements PCLM as the core model architecture.
  • Multi-model integration: Integrates three distinct protein language models that characterize sequence, structural features, and functional attributes.
  • Attention-based alignment: Uses an attention-based alignment mechanism to generate a comprehensive joint representation of proteins.
  • Training strategy: PCLM is pre-trained on a large-scale dataset of IDR sequences and fine-tuned on functionally annotated sequences for molecular function prediction.
  • Addressing prior limitations: Extends predictive coverage beyond methods that predominantly focus on entropic chain functions to the other five categories of disordered functions.
  • Performance evaluation: Demonstrated high-quality predictive performance across five categories of functional and multi-functional residues.

Scientific Applications:

  • Drug target discovery and rational drug design: Supports identification of disordered protein functions relevant to drug target discovery and rational drug design.
  • Computational functional annotation of IDPs/IDRs: Enables research-scale prediction of molecular functions of intrinsically disordered proteins and regions for downstream biological studies.

Methodology:

DMFpred implements the Protein Cubic Language Model (PCLM), which integrates three protein language models (sequence, structural, functional) via an attention-based alignment mechanism; PCLM was pre-trained on a large-scale dataset of IDR sequences and fine-tuned on functionally annotated sequences for molecular function prediction.

Topics

Details

License:
Not licensed
Cost:
Free of charge
Tool Type:
web application
Operating Systems:
Mac, Linux, Windows
Added:
1/28/2023
Last Updated:
11/24/2024

Operations

Publications

Pang Y, Liu B. DMFpred: Predicting protein disorder molecular functions based on protein cubic language model. PLOS Computational Biology. 2022;18(10):e1010668. doi:10.1371/journal.pcbi.1010668. PMID:36315580. PMCID:PMC9674156.

PMID: 36315580
PMCID: PMC9674156
Funding: - National Natural Science Foundation of China: 62271049 - Beijing Natural Science Foundation: JQ19019