Gilda
Gilda grounds named entities in biomedical literature by scored string matching and machine-learned disambiguation to map mentions to identifiers in biomedical ontologies.
Key Features:
- Scored string matching algorithm: Employs a scored string matching algorithm to identify names and synonyms across ontology entries.
- Machine-learned disambiguation models: Uses machine-learned models that leverage contextual information from surrounding text to resolve ambiguous entity names.
- Species-prioritization: Applies species-prioritization to refine disambiguation based on biological context.
- Ontology coverage: Grounds entities across biomedical ontologies covering genes, proteins (including families and complexes), small molecules, biological processes, and diseases.
- Identifier selection: Selects the correct identifiers when multiple potential matches exist for a given string.
Scientific Applications:
- Biomedical text mining: Automates grounding of entity mentions to support large-scale text mining and annotation of literature.
- Data extraction and curation: Facilitates reliable extraction and annotation of entities from high volumes of publications where manual curation is impractical.
- Normalization for analysis: Provides normalized identifiers for genes, proteins, small molecules, biological processes, and diseases to enable downstream data integration and analysis.
Methodology:
Implements a scored string matching algorithm combined with machine-learned disambiguation models that leverage surrounding text and species-prioritization to map mentions to ontology identifiers.
Topics
Details
- License:
- BSD-2-Clause
- Cost:
- Free of charge
- Tool Type:
- command-line tool
- Operating Systems:
- Mac, Linux, Windows
- Programming Languages:
- Python
- Added:
- 3/1/2022
- Last Updated:
- 3/1/2022
Operations
Publications
Gyori BM, Hoyt CT, Steppi A. Gilda: biomedical entity text normalization with machine-learned disambiguation as a service. Unknown Journal. 2021. doi:10.1101/2021.09.10.459803.