ProfhEX

ProfhEX predicts early-stage off-target liabilities and toxicity profiles of small molecules to support liability profiling in drug discovery.


Key Features:

  • OECD-compliant ensemble: An ensemble of 46 machine learning models developed and validated according to OECD principles for liability prediction.
  • Liability coverage: Profiles seven liability groups: cardiovascular, central nervous system, gastrointestinal, endocrine, renal, pulmonary, and immune system toxicities.
  • Training data: Models trained on experimental affinity data from public and commercial sources encompassing 289,202 activity data points for 210,116 unique compounds across 46 targets with dataset sizes ranging from 819 to 18,896.
  • Algorithms: Uses gradient boosting and random forest algorithms to select champion models within the ensemble.
  • Validation methods: Employs internal validation techniques including cross-validation, bootstrap methods, and y‑scrambling, alongside external validation in accordance with OECD guidelines.
  • Predictive performance: Reported average Pearson correlation r = 0.84 (SD = 0.05), R² = 0.68 (SD = 0.1), and RMSE = 0.69 (SD = 0.08).
  • Hit-detection metrics: Demonstrates average enrichment factor at 5% = 13.1 (SD = 4.5) and AUC = 0.92 (SD = 0.05).
  • Benchmarking and integration: Benchmarked against existing tools for large-scale liability profiling and includes integration of complementary modeling approaches such as structure-based and pharmacophore-based models.

Scientific Applications:

  • Off-target liability profiling: Predicts binding-related liabilities across seven biological systems to inform compound risk assessment.
  • Early-stage candidate prioritization: Ranks small molecules by predicted affinity and liability to guide selection in lead optimization.
  • Virtual screening enrichment: Supports hit detection and prioritization in screening campaigns using reported EF and AUC performance.
  • Validated model support: Provides OECD-compliant model outputs suitable for benchmarking and validation workflows relevant to regulatory and preclinical assessment.

Methodology:

Models were trained on experimental affinity data from public and commercial sources (289,202 data points for 210,116 compounds across 46 targets); gradient boosting and random forest algorithms were used to nominate champion models within an ensemble of 46 machine learning models; validation followed OECD principles using cross‑validation, bootstrap, y‑scrambling and external validation, and models were benchmarked against existing tools.

Topics

Details

Cost:
Free of charge
Tool Type:
web application
Operating Systems:
Mac, Linux, Windows
Added:
1/23/2024
Last Updated:
11/24/2024

Operations

Publications

Lunghini F, Fava A, Pisapia V, Sacco F, Iaconis D, Beccari AR. ProfhEX: AI-based platform for small molecules liability profiling. Journal of Cheminformatics. 2023;15(1). doi:10.1186/s13321-023-00728-6. PMID:37296454. PMCID:PMC10251600.