DMFpred
DMFpred predicts molecular functions of intrinsically disordered proteins and regions (IDPs/IDRs) at the residue level to identify functionally annotated residues relevant to drug target discovery and rational drug design.
Key Features:
- Function scope: Predicts six disordered molecular function types: entropic chain, assembler, scavenger, effector, display site, and chaperone.
- Residue-level and multi-functional prediction: Provides residue-level predictions and supports identification of functional and multi-functional residues across five evaluation categories.
- Protein Cubic Language Model (PCLM): Implements PCLM as the core model architecture.
- Multi-model integration: Integrates three distinct protein language models that characterize sequence, structural features, and functional attributes.
- Attention-based alignment: Uses an attention-based alignment mechanism to generate a comprehensive joint representation of proteins.
- Training strategy: PCLM is pre-trained on a large-scale dataset of IDR sequences and fine-tuned on functionally annotated sequences for molecular function prediction.
- Addressing prior limitations: Extends predictive coverage beyond methods that predominantly focus on entropic chain functions to the other five categories of disordered functions.
- Performance evaluation: Demonstrated high-quality predictive performance across five categories of functional and multi-functional residues.
Scientific Applications:
- Drug target discovery and rational drug design: Supports identification of disordered protein functions relevant to drug target discovery and rational drug design.
- Computational functional annotation of IDPs/IDRs: Enables research-scale prediction of molecular functions of intrinsically disordered proteins and regions for downstream biological studies.
Methodology:
DMFpred implements the Protein Cubic Language Model (PCLM), which integrates three protein language models (sequence, structural, functional) via an attention-based alignment mechanism; PCLM was pre-trained on a large-scale dataset of IDR sequences and fine-tuned on functionally annotated sequences for molecular function prediction.
Topics
Details
- License:
- Not licensed
- Cost:
- Free of charge
- Tool Type:
- web application
- Operating Systems:
- Mac, Linux, Windows
- Added:
- 1/28/2023
- Last Updated:
- 11/24/2024
Operations
Publications
Pang Y, Liu B. DMFpred: Predicting protein disorder molecular functions based on protein cubic language model. PLOS Computational Biology. 2022;18(10):e1010668. doi:10.1371/journal.pcbi.1010668. PMID:36315580. PMCID:PMC9674156.