BioRED

BioRED provides a document-level biomedical relation extraction dataset that annotates genes/proteins, diseases, and chemicals across diverse relation pairs and labels each relation for novelty to support development and evaluation of automated RE and NER systems.


Key Features:

  • Multiple Entity Types: Includes annotations for genes/proteins, diseases, and chemicals.
  • Diverse Relation Pairs: Contains various relation pairs, including gene-disease and chemical-chemical interactions.
  • Document-Level Scope: Operates at the document level to capture context-rich relationships beyond single sentences.
  • Novelty Annotation: Labels each relation as a novel finding or previously known background knowledge.

Scientific Applications:

  • Relation Extraction development: Enables training and benchmarking of RE systems, including models based on Bidirectional Encoder Representations from Transformers (BERT).
  • Named Entity Recognition benchmarking: Serves as a resource for evaluating NER performance on genes/proteins, diseases, and chemicals.
  • Novelty analysis in biomedical text: Facilitates research into distinguishing novel findings from background knowledge in biomedical literature.

Methodology:

Constructed after a review of existing NER and RE datasets, assembled from 600 PubMed abstracts, and evaluated via benchmarking experiments with state-of-the-art methods that reported high NER performance while highlighting challenges in extracting novel relations.

Topics

Details

License:
Not licensed
Cost:
Free of charge
Tool Type:
web application
Operating Systems:
Mac, Linux, Windows
Added:
9/26/2022
Last Updated:
9/26/2022

Operations

Publications

Luo L, Lai P, Wei C, Arighi CN, Lu Z. BioRED: a rich biomedical relation extraction dataset. Briefings in Bioinformatics. 2022;23(5). doi:10.1093/bib/bbac282. PMID:35849818. PMCID:PMC9487702.

PMID: 35849818
Funding: - National Institutes of Health: 2U24HG007822-08