DOUBLER

DOUBLER learns unified representations of biological entities by integrating a knowledge graph with structured annotations and free-text documents to predict protein–disease associations.


Key Features:

  • Unified Representation Learning: Integrates heterogeneous data sources into a single representation space for proteins, diseases, and associated documents.
  • Knowledge Graph Utilization: Employs a knowledge graph to interlink biological entities and contextualize relationships between proteins and diseases.
  • Multiple Data Modalities: Explicitly incorporates structured annotations and free-text descriptions alongside graph data into the learning process.
  • Consistency Enforcement: Incorporates an objective that enforces consistency across representations derived from different modalities within the learning algorithm.
  • Link Prediction Performance: Improves link prediction for protein–disease associations relative to state-of-the-art link prediction algorithms when informative free text is available.

Scientific Applications:

  • Protein–Disease Association Prediction: Predicts novel protein–disease associations to identify potential links between proteins and diseases.
  • Drug Target Prioritization: Prioritizes protein targets for therapeutic investigation based on predicted disease associations.

Methodology:

Learns representations by integrating a knowledge graph with additional modalities (structured annotations and free text) and explicitly incorporates representation consistency into the learning objective for link prediction of protein–disease associations.

Topics

Details

Tool Type:
command-line tool
Programming Languages:
Python
Added:
1/18/2021
Last Updated:
3/1/2021

Operations

Publications

Sztyler T, Malone B. DOUBLER: Unified Representation Learning of Biological Entities and Documents for Predicting Protein–Disease Relationships. Unknown Journal. 2020. doi:10.1101/2020.10.27.357202.