SmuDGE
SmuDGE generates semantic embeddings of phenotypes to prioritize disease-associated genes by comparing phenotype sets and leveraging interaction-network-derived indirect associations.
Key Features:
- Semantic Similarity Measure: Utilizes a trainable semantic similarity measure to compare sets of phenotypes between entities such as disease-gene or disease-patient pairs.
- Feature Learning and Vector Representations: Produces vector-based phenotype representations that capture relationships from direct phenotype annotations and learned features.
- Exploitation of Interaction Networks: Exploits background knowledge from interaction networks, including protein-protein interactions, to derive phenotype representations for indirectly associated entities.
- Enhanced Coverage and Performance: Extends coverage to all genes within a connected interaction network and can match or outperform traditional semantic similarity methods in phenotype-based disease gene prioritization tasks.
Scientific Applications:
- Disease Gene Prioritization: Ranks and identifies candidate disease-associated genes, including cases with sparse or missing direct phenotype annotations in human and model organism studies.
- Phenotype-Based Comparisons: Compares phenotypic profiles between diseases, genes, or patients to support investigations of disease mechanisms and translational or personalized-medicine analyses.
Methodology:
SmuDGE uses Python scripts (e.g., smudge_E.py) to generate corpora and feature vectors by sampling nodes' surrounding environments in interaction networks, extracts phenotypic information from the PhenomNet ontology including associated and superclass terms, trains a semantic similarity measure, and creates vector-based semantic embeddings that capture direct and indirect phenotype associations.
Topics
Collections
Details
- Tool Type:
- command-line tool
- Programming Languages:
- Python, C++, Groovy
- Added:
- 1/20/2021
- Last Updated:
- 5/20/2021
Operations
Publications
Alshahrani M, Hoehndorf R. Semantic Disease Gene Embeddings (SmuDGE): phenotype-based disease gene prioritization without phenotypes. Bioinformatics. 2018;34(17):i901-i907. doi:10.1093/bioinformatics/bty559. PMID:30423077. PMCID:PMC6129260.