PheneBank
PheneBank extracts and validates human phenotype–disease associations from Medline using machine learning–based text mining to integrate literature-derived phenotypic annotations into ontologies such as the Human Phenotype Ontology (HPO).
Key Features:
- Text Mining Capabilities: Employs machine learning algorithms to identify phenotype concepts in Medline using an expert-annotated rare disease corpus from the PMC Text Mining subset.
- Integration with Existing Ontologies: Produces literature-based phenotype annotations that can be mapped to the Human Phenotype Ontology (HPO) and other coding systems.
- Validation Against Gold Standards: Evaluates performance using a gold-standard corpus of rare disease sentences and cross-references extracted associations with the Monarch Initiative.
Scientific Applications:
- Phenotype–disease association discovery: Automates extraction of candidate phenotype–disease links from literature to support discovery of novel associations.
- Phenotype database curation: Supplements and expands phenotype databases by providing literature-derived annotations for integration into ontologies and coding systems.
- Rare disease research: Enhances aggregation and validation of detailed phenotypic information pertinent to rare disease analyses and hypothesis generation.
Methodology:
Machine learning–based text mining models trained on expert-annotated rare disease datasets from the PMC Text Mining subset and Medline perform contextual analysis to identify phenotype mentions; performance is evaluated against a gold-standard corpus and cross-referenced with the Monarch Initiative.
Topics
Collections
Details
- License:
- MIT
- Tool Type:
- web application
- Programming Languages:
- Python
- Added:
- 1/17/2022
- Last Updated:
- 1/17/2022
Operations
Publications
Pilehvar MT, Bernard A, Smedley D, Collier N. PheneBank: a literature-based database of phenotypes. Bioinformatics. 2021;38(4):1179-1180. doi:10.1093/bioinformatics/btab740. PMID:34788791. PMCID:PMC8796364.