PhenoPred

PhenoPred predicts novel human gene–disease associations by integrating known gene–disease associations, protein–protein interaction networks, molecular functional annotations, and protein sequence features using supervised machine learning (support vector machines).


Key Features:

  • Data integration: Integrates known gene–disease associations, protein–protein interaction (PPI) networks, molecular-level functional annotations, and protein sequence information.
  • Machine learning: Employs a supervised machine learning approach with support vector machines (SVMs) as the core algorithmic component.
  • Network-based mapping: Maps each gene or protein onto spaces defined by disease terms and functional annotations using proximity within the human PPI network.
  • Feature encoding: Encodes sequence characteristics, functional attributes, physicochemical properties, and predicted structural features such as secondary structure and flexibility.
  • Robustness to noisy data: Handles noisy and incomplete experimental data.
  • Evolving ontology support: Incorporates an evolving ontology of diseases in its analysis.
  • Multi-disease association detection: Identifies candidate genes associated with multiple disease terms simultaneously.
  • Prioritization and prediction: Prioritizes genes by likelihood of association and identifies potential diseases linked to a given gene.

Scientific Applications:

  • Candidate gene prioritization: Ranks and prioritizes candidate genes based on predicted association likelihood with diseases.
  • Exploration of complex relationships: Facilitates exploration of complex gene–disease relationships across multiple disease terms.
  • Disease prediction for genes: Identifies potential diseases linked to a specific gene.
  • Support for translational research: Provides evidence-based predictions to aid personalized medicine and the discovery of novel therapeutic targets.

Methodology:

Integrates known gene–disease associations, human PPI networks, molecular functional annotations, and protein sequence information; maps genes/proteins into spaces defined by disease terms and functional annotations using PPI proximity; encodes sequence characteristics, functional attributes, physicochemical properties, and predicted structural features (secondary structure and flexibility); and applies supervised machine learning using support vector machines while handling noisy and incomplete experimental data and incorporating an evolving disease ontology.

Topics

Details

Tool Type:
api
Operating Systems:
Linux, Windows, Mac
Added:
8/3/2017
Last Updated:
11/25/2024

Operations

Publications

Radivojac P, Peng K, Clark WT, Peters BJ, Mohan A, Boyle SM, Mooney SD. An integrated approach to inferring gene–disease associations in humans. Proteins: Structure, Function, and Bioinformatics. 2008;72(3):1030-1037. doi:10.1002/prot.21989. PMID:18300252. PMCID:PMC2824611.

PMID: 18300252
PMCID: PMC2824611
Funding: - NSF: DBI-0644017 - NIH: K22LM009135, P01AG018397

Documentation

Links