BRONCO

BRONCO provides a manually annotated corpus for information extraction from German oncological discharge summaries, supporting development and evaluation of named entity recognition, entity normalization, and negation and speculation detection.


Key Features:

  • Data composition: The corpus comprises 200 manually deidentified oncological discharge summaries totaling 11,434 sentences and 89,942 tokens.
  • Annotations: The dataset contains 11,124 annotations for medical entities (including diagnoses, treatments, and medications) and 3,118 annotations for attributes such as negation and speculation.
  • Annotation quality: A structured, quality-controlled annotation process was carried out involving two groups of medical experts to ensure consistency and comprehensiveness.
  • Data structure: Seventy-five percent of the corpus is published as a set of shuffled sentences while 25% is reserved as a held-out dataset for evaluation.
  • Baseline results: Baseline evaluations using state-of-the-art techniques on the held-out dataset report F1-scores ranging from 0.10 to 0.90 across different tasks.

Scientific Applications:

  • Information extraction from German medical texts: The corpus supports development and benchmarking of IE systems on German oncological discharge summaries.
  • Named entity recognition (NER): Annotations enable training and evaluation of NER models for diagnoses, treatments, and medications.
  • Entity normalization: Annotated entity mentions support mapping to normalized concepts or terminologies.
  • Negation detection: Negation annotations enable development and testing of negation detection methods.
  • Speculation detection: Speculation annotations enable development and testing of speculation/uncertainty detection methods.

Methodology:

The corpus was created by manual deidentification of 200 discharge summaries followed by structured annotation by two groups of medical experts; the dataset is split into 75% shuffled sentences and a 25% held-out set, and baseline evaluations with state-of-the-art techniques yielded F1-scores of 0.10–0.90.

Topics

Details

Tool Type:
database
Added:
6/14/2021
Last Updated:
8/18/2021

Operations

Publications

Kittner M, Lamping M, Rieke DT, Götze J, Bajwa B, Jelas I, Rüter G, Hautow H, Sänger M, Habibi M, Zettwitz M, Bortoli Td, Ostermann L, Ševa J, Starlinger J, Kohlbacher O, Malek NP, Keilholz U, Leser U. Annotation and initial evaluation of a large annotated German oncological corpus. JAMIA Open. 2021;4(2). doi:10.1093/jamiaopen/ooab025. PMID:33898938. PMCID:PMC8054032.

PMID: 33898938
PMCID: PMC8054032
Funding: - German Bundesministerium für Bildung und Forschung: 031L0023B, 031L0030B - Deutsche Forschungsgemeinschaft: LE1428/1-1