GeneralizedDTA

GeneralizedDTA predicts drug-target binding affinity by combining self-supervised pre-training on protein amino acid sequences and drug molecular graphs with multi-task learning and a dual adaptation mechanism to improve generalization to unknown drugs and out-of-distribution compounds.


Key Features:

  • Pre-training for structural information: Employs self-supervised pre-training tasks on amino acid sequences and molecular graphs to capture detailed structural information of proteins and drug compounds.
  • Mitigation of encoding variance: Reduces high variance associated with deep neural network-based encoding by enriching learned feature representations via pre-training.
  • Multi-task learning framework: Uses a multi-task learning framework with a dual adaptation mechanism to narrow the task gap between pre-training and DTA prediction and to prevent overfitting.
  • Handling out-of-distribution problems: Addresses out-of-distribution issues to improve predictive performance on novel or unrepresented drugs in labeled datasets.

Scientific Applications:

  • Drug discovery: Predicts binding affinities for unknown drugs to accelerate identification of potential therapeutic candidates.
  • Model generalization: Enhances generalization across diverse datasets to enable DTA prediction with limited labeled data.

Methodology:

Performs self-supervised pre-training on large-scale unlabeled data of amino acid sequences and molecular graphs, followed by a multi-task learning phase with a dual adaptation mechanism to adapt pre-trained representations for DTA prediction.

Topics

Details

License:
Not licensed
Cost:
Free of charge
Tool Type:
command-line tool
Programming Languages:
Python
Added:
11/4/2022
Last Updated:
11/24/2024

Operations

Publications

Lin S, Shi C, Chen J. GeneralizedDTA: combining pre-training and multi-task learning to predict drug-target binding affinity for unknown drug discovery. BMC Bioinformatics. 2022;23(1). doi:10.1186/s12859-022-04905-6. PMID:36071406. PMCID:PMC9449940.

PMID: 36071406
PMCID: PMC9449940
Funding: - National Key Research and Development Program of China: 2020YFB2104402 - Beijing Municipal Natural Science Foundation: 4222022