Gilda

Gilda grounds named entities in biomedical literature by scored string matching and machine-learned disambiguation to map mentions to identifiers in biomedical ontologies.


Key Features:

  • Scored string matching algorithm: Employs a scored string matching algorithm to identify names and synonyms across ontology entries.
  • Machine-learned disambiguation models: Uses machine-learned models that leverage contextual information from surrounding text to resolve ambiguous entity names.
  • Species-prioritization: Applies species-prioritization to refine disambiguation based on biological context.
  • Ontology coverage: Grounds entities across biomedical ontologies covering genes, proteins (including families and complexes), small molecules, biological processes, and diseases.
  • Identifier selection: Selects the correct identifiers when multiple potential matches exist for a given string.

Scientific Applications:

  • Biomedical text mining: Automates grounding of entity mentions to support large-scale text mining and annotation of literature.
  • Data extraction and curation: Facilitates reliable extraction and annotation of entities from high volumes of publications where manual curation is impractical.
  • Normalization for analysis: Provides normalized identifiers for genes, proteins, small molecules, biological processes, and diseases to enable downstream data integration and analysis.

Methodology:

Implements a scored string matching algorithm combined with machine-learned disambiguation models that leverage surrounding text and species-prioritization to map mentions to ontology identifiers.

Topics

Details

License:
BSD-2-Clause
Cost:
Free of charge
Tool Type:
command-line tool
Operating Systems:
Mac, Linux, Windows
Programming Languages:
Python
Added:
3/1/2022
Last Updated:
3/1/2022

Operations

Publications

Gyori BM, Hoyt CT, Steppi A. Gilda: biomedical entity text normalization with machine-learned disambiguation as a service. Unknown Journal. 2021. doi:10.1101/2021.09.10.459803.

Downloads