CoKE
CoKE extracts and annotates drug–protein relationships from the COVID Open Research Dataset (CORD-19) to support identification of therapeutic candidates for SARS-CoV-2.
Key Features:
- Data Source: Uses the COVID Open Research Dataset (CORD-19) as the literature corpus for mining drug and protein mentions.
- Ontological Tagging and Entity Resolution: Applies SciBiteAI ontological tagging and resolves entities to UniProt for proteins and DrugBank for drugs.
- Relationship Identification Algorithm: Employs a custom algorithm that detects co-occurrences of protein and drug terms and assigns confidence scores to identified pairs.
- Processed Results: Has processed CORD-19 to identify approximately 3,000 drug–protein pairs involving 29 unique proteins and over 500 investigational, experimental, and approved drugs, including compounds in clinical trials.
Scientific Applications:
- Drug repurposing for SARS-CoV-2: Prioritizes drug–protein associations to support identification of candidate therapeutics against SARS-CoV-2.
- Literature-scale knowledge extraction: Condenses the CORD-19 corpus into curated drug–target relationships for downstream analysis and hypothesis generation.
Methodology:
Automated extraction from CORD-19 using SciBiteAI ontological tagging; mapping of extracted entities to UniProt and DrugBank identifiers; detection of co-occurrences between drug and protein terms via a custom algorithm that calculates confidence scores for each pair.
Topics
Collections
Details
- Tool Type:
- command-line tool, web application
- Added:
- 1/18/2021
- Last Updated:
- 2/14/2021
Operations
Publications
Korn D, Pervitsky V, Bobrowski T, Alves V, Schmitt C, Bizon C, Baker N, Chirkova R, Cherkasov A, Muratov E, Tropsha A. COVID-19 Knowledge Extractor (COKE): A Tool and a Web Portal to Extract Drug - Target Protein Associations from the CORD-19 Corpus of Scientific Publications on COVID-19. Unknown Journal. 2020. doi:10.26434/chemrxiv.13289222.v1.