ACELA
ACELA streamlines the creation of comprehensive, high-quality annotated corpora for named entity (NE) recognition by combining probabilistic NE tagging with an iterative human-in-the-loop annotation process to reduce manual annotation effort.
Key Features:
- Active Learning-Like Framework: Employs an active learning–inspired approach that systematically annotates all occurrences of target named entities within a corpus to avoid sampling bias.
- Iterative and Interactive Process: Facilitates repeated interaction between human annotators and a probabilistic NE tagger to refine annotations over successive rounds.
- Reduction in Human Effort: Simulates annotation and leverages the probabilistic tagger to minimize the number of sentences requiring manual examination, especially for sparse target named entities.
- Probabilistic Tagging: Uses a probabilistic model to predict named entities within text and guide annotator focus.
- Iterative Refinement: Updates the probabilistic model based on human feedback to improve predictive accuracy across iterations.
- Comprehensive Coverage: Ensures annotation of all instances of target entities to eliminate biases associated with selective sampling.
Scientific Applications:
- Biomedical informatics: Enables creation of gold-standard annotated datasets for NE recognition used in biomedical text mining and extraction tasks.
- Genomics: Supports extraction and annotation of genomic entities from large text corpora to improve entity recognition models in genomics research.
- Information extraction from large corpora: Applies to domains requiring detailed entity extraction and high-quality annotated corpora for training NE recognizers.
Methodology:
Probabilistic tagging of text using a probabilistic NE model; iterative refinement of the model based on human annotator feedback; systematic annotation to ensure comprehensive coverage of target entities.
Topics
Details
- Tool Type:
- web application
- Operating Systems:
- Linux, Windows, Mac
- Added:
- 5/1/2017
- Last Updated:
- 11/25/2024
Operations
Publications
Tsuruoka Y, Tsujii J, Ananiadou S. Accelerating the annotation of sparse named entities by dynamic sentence selection. BMC Bioinformatics. 2008;9(S11). doi:10.1186/1471-2105-9-s11-s8. PMID:19025694. PMCID:PMC2586757.