NEREL-BIO
NEREL-BIO provides nested named entity annotations and bilingual biomedical corpora to support cross-domain and cross-language named entity recognition research.
Key Features:
- Nested Named Entity Recognition: Identifies and annotates nested named entities in biomedical texts, including shorter entities embedded within longer ones that can span boundaries.
- Cross-Domain and Cross-Language Transfer: Supports domain transfer experiments from NEREL → NEREL-BIO and provides parallel English–Russian annotations of PubMed abstracts for cross-language research.
- Extensive Corpus: Contains over 700 Russian and 100 English biomedical PubMed abstracts, with each English annotation having a corresponding Russian counterpart.
- Annotation Scheme: Employs an annotation scheme that integrates general-domain entities and specialized biomedical entity types.
Scientific Applications:
- NER Model Development: Enables training and evaluation of NER models, including transformer architectures and machine reading comprehension frameworks, on nested-entity and bilingual corpora.
- Cross-Domain and Cross-Language Studies: Facilitates computational linguistics and biomedical text mining studies requiring adaptation across domains or analysis of multilingual datasets.
Methodology:
Explicit methods include nested named entity annotation, domain transfer experiments (NEREL → NEREL-BIO), parallel English and Russian PubMed abstract annotations, and evaluation with transformer-based and machine reading comprehension NER models.
Topics
Details
- Tool Type:
- database
- Added:
- 9/25/2023
- Last Updated:
- 11/24/2024
Operations
Data Inputs & Outputs
Named-entity and concept recognition
Publications
Loukachevitch N, Manandhar S, Baral E, Rozhkov I, Braslavski P, Ivanov V, Batura T, Tutubalina E. NEREL-BIO: a dataset of biomedical abstracts annotated with nested named entities. Bioinformatics. 2023;39(4). doi:10.1093/bioinformatics/btad161. PMID:37004189. PMCID:PMC10129873.