ACELA

ACELA streamlines the creation of comprehensive, high-quality annotated corpora for named entity (NE) recognition by combining probabilistic NE tagging with an iterative human-in-the-loop annotation process to reduce manual annotation effort.


Key Features:

  • Active Learning-Like Framework: Employs an active learning–inspired approach that systematically annotates all occurrences of target named entities within a corpus to avoid sampling bias.
  • Iterative and Interactive Process: Facilitates repeated interaction between human annotators and a probabilistic NE tagger to refine annotations over successive rounds.
  • Reduction in Human Effort: Simulates annotation and leverages the probabilistic tagger to minimize the number of sentences requiring manual examination, especially for sparse target named entities.
  • Probabilistic Tagging: Uses a probabilistic model to predict named entities within text and guide annotator focus.
  • Iterative Refinement: Updates the probabilistic model based on human feedback to improve predictive accuracy across iterations.
  • Comprehensive Coverage: Ensures annotation of all instances of target entities to eliminate biases associated with selective sampling.

Scientific Applications:

  • Biomedical informatics: Enables creation of gold-standard annotated datasets for NE recognition used in biomedical text mining and extraction tasks.
  • Genomics: Supports extraction and annotation of genomic entities from large text corpora to improve entity recognition models in genomics research.
  • Information extraction from large corpora: Applies to domains requiring detailed entity extraction and high-quality annotated corpora for training NE recognizers.

Methodology:

Probabilistic tagging of text using a probabilistic NE model; iterative refinement of the model based on human annotator feedback; systematic annotation to ensure comprehensive coverage of target entities.

Topics

Details

Tool Type:
web application
Operating Systems:
Linux, Windows, Mac
Added:
5/1/2017
Last Updated:
11/25/2024

Operations

Publications

Tsuruoka Y, Tsujii J, Ananiadou S. Accelerating the annotation of sparse named entities by dynamic sentence selection. BMC Bioinformatics. 2008;9(S11). doi:10.1186/1471-2105-9-s11-s8. PMID:19025694. PMCID:PMC2586757.

Documentation