MarkerGenie
MarkerGenie extracts biomedical entity relations from unstructured textual sources such as scientific articles and dialogues using natural language processing to produce structured relation data for biomedical marker discovery.
Key Features:
- Automatic Training Data Labeling: Incorporates an automatic training data labeling mechanism to generate labeled datasets in the biomedical domain where labeled data are scarce.
- Rule-Based Biological Terminology Cleaning: Applies rule-based cleaning methods to address ambiguities and inconsistencies in biological terminology.
- Advanced NLP Model for Relation Prediction: Integrates an NLP model capable of binary associative and multi-relation prediction between biomedical entities.
- Unstructured-to-Structured Transformation: Transforms unstructured text from scientific articles and dialogues into structured information suitable for downstream analysis.
Scientific Applications:
- Biomarker Discovery: Automates extraction of relations among biomedical entities to support marker discovery studies.
- Literature Survey and Curation: Enables comprehensive surveys of the literature and helps reduce biases associated with manual curation.
- Research Orientation and R&D Prioritization: Rapidly processes large volumes of text to assist researchers in orienting research and development efforts.
Methodology:
Uses automatic training data labeling, rule-based biological terminology cleaning, and advanced NLP techniques for binary associative and multi-relation prediction; performance demonstrated on benchmark datasets and case studies.
Topics
Details
- License:
- CC-BY-4.0
- Cost:
- Free of charge
- Tool Type:
- web application
- Operating Systems:
- Mac, Linux, Windows
- Added:
- 1/2/2024
- Last Updated:
- 11/24/2024
Operations
Publications
Gu W, Yang X, Yang M, Han K, Pan W, Zhu Z. MarkerGenie: an NLP-enabled text-mining system for biomedical entity relation extraction. Bioinformatics Advances. 2022;2(1). doi:10.1093/bioadv/vbac035. PMID:36699388. PMCID:PMC9710573.
PMID: 36699388
PMCID: PMC9710573
Funding: - National Natural Science Foundation of China: 61871272 and 61911530218
- Guangdong Provincial Key Laboratory: 2020B121201001
- Shenzhen Fundamental Research Program: JCYJ20190808173617147
- BGIShenzhen: BGIRSZ20200002