Gimli

Gimli performs biomedical named-entity recognition to identify and extract biomedical names from scientific literature for biomedical information extraction.


Key Features:

  • Comprehensive feature set: Implements orthographic, morphological, linguistic-based, conjunctions, and dictionary-based features for NER.
  • Model combination methodology: Provides a method to combine different trained models to leverage multiple model strengths.
  • High performance: Reports F-measure values of 87.17% on the GENETAG corpus and 72.23% on the JNLPBA corpus.
  • Integration and extensibility: Exposes a library interface and allows extension or adaptation of its functionalities.

Scientific Applications:

  • Biomedical named-entity recognition: Identification and extraction of gene, protein, and other biomedical names from scientific text.
  • Biomedical information extraction and text mining: Support for bioinformatics and computational biology projects requiring accurate NER as an input to downstream extraction tasks.

Methodology:

Implements multiple feature types (orthographic, morphological, linguistic-based, conjunctions, dictionary-based) and a model combination approach; supports pre-trained models and user-defined training processes.

Topics

Details

License:
NCSA
Tool Type:
desktop application
Operating Systems:
Linux, Windows, Mac
Programming Languages:
Java
Added:
5/6/2018
Last Updated:
12/10/2018

Operations

Publications

Campos D, Matos S, Oliveira JL. Gimli: open source and high-performance biomedical name recognition. BMC Bioinformatics. 2013;14(1). doi:10.1186/1471-2105-14-54. PMID:23413997. PMCID:PMC3651325.

Documentation