CoKE

CoKE extracts and annotates drug–protein relationships from the COVID Open Research Dataset (CORD-19) to support identification of therapeutic candidates for SARS-CoV-2.


Key Features:

  • Data Source: Uses the COVID Open Research Dataset (CORD-19) as the literature corpus for mining drug and protein mentions.
  • Ontological Tagging and Entity Resolution: Applies SciBiteAI ontological tagging and resolves entities to UniProt for proteins and DrugBank for drugs.
  • Relationship Identification Algorithm: Employs a custom algorithm that detects co-occurrences of protein and drug terms and assigns confidence scores to identified pairs.
  • Processed Results: Has processed CORD-19 to identify approximately 3,000 drug–protein pairs involving 29 unique proteins and over 500 investigational, experimental, and approved drugs, including compounds in clinical trials.

Scientific Applications:

  • Drug repurposing for SARS-CoV-2: Prioritizes drug–protein associations to support identification of candidate therapeutics against SARS-CoV-2.
  • Literature-scale knowledge extraction: Condenses the CORD-19 corpus into curated drug–target relationships for downstream analysis and hypothesis generation.

Methodology:

Automated extraction from CORD-19 using SciBiteAI ontological tagging; mapping of extracted entities to UniProt and DrugBank identifiers; detection of co-occurrences between drug and protein terms via a custom algorithm that calculates confidence scores for each pair.

Topics

Collections

Details

Tool Type:
command-line tool, web application
Added:
1/18/2021
Last Updated:
2/14/2021

Operations

Publications

Korn D, Pervitsky V, Bobrowski T, Alves V, Schmitt C, Bizon C, Baker N, Chirkova R, Cherkasov A, Muratov E, Tropsha A. COVID-19 Knowledge Extractor (COKE): A Tool and a Web Portal to Extract Drug - Target Protein Associations from the CORD-19 Corpus of Scientific Publications on COVID-19. Unknown Journal. 2020. doi:10.26434/chemrxiv.13289222.v1.

Links