PrePCI

PrePCI predicts interactions between chemical compounds and human proteins across a proteome-wide set of 19,797 human proteins and 6.8 million chemical compounds (yielding over 5 billion predicted interactions) and provides likelihood ratios quantifying binding likelihood based on sequence, structural, and chemical similarity using the AlphaFold Protein Structure Database, Protein Data Bank templates, LT-scanner, the Tanimoto coefficient, and Naive Bayesian statistics.


Key Features:

  • Proteome-wide structural models: Integrates traditional protein modeling techniques with structures from the AlphaFold Protein Structure Database to generate proteome-scale models.
  • Database scale: Contains over 5 billion predicted interactions involving 6.8 million chemical compounds and 19,797 human proteins.
  • PDB template mapping: Establishes mappings between template proteins in the Protein Data Bank known to bind specific compounds and query proteins in the model database.
  • Sequence- and structure-based similarity: Computes sequence- and structure-based similarity metrics between template (T) and query (Q) proteins to assess potential transfer of ligand binding.
  • Threshold-based inference: Infers compound C binding to query protein Q when similarity metrics exceed predefined threshold values.
  • LT-scanner structural scoring: Uses the LT-scanner scoring function to assess how well a compound fits within a predicted binding site for structure-based predictions.
  • Chemical similarity: Evaluates chemical similarity between compounds using the Tanimoto coefficient to identify related small molecules.
  • Likelihood ratios and statistical integration: Derives likelihood ratios from machine learning algorithms and integrates structural and chemical evidence using Naive Bayesian statistics.

Scientific Applications:

  • Lead discovery: Supports identification and prioritization of novel small-molecule leads by predicting protein–compound interactions.
  • Drug mechanism elucidation: Aids interpretation of drug–target interactions by providing predicted binding partners and likelihood ratios.
  • Biological function annotation: Enables inference of protein biological roles from predicted interaction profiles with small molecules.

Methodology:

PrePCI builds proteome-wide structural models from traditional modeling and the AlphaFold Protein Structure Database, computes sequence- and structure-based similarity metrics between Protein Data Bank template proteins (T) that bind compounds (C) and query proteins (Q), applies threshold-based inference, scores structural fits with LT-scanner, assesses chemical similarity with the Tanimoto coefficient, derives likelihood ratios via machine learning, and integrates structural and chemical evidence using Naive Bayesian statistics.

Topics

Details

Cost:
Free of charge
Tool Type:
web application
Operating Systems:
Mac, Linux, Windows
Added:
3/27/2023
Last Updated:
11/24/2024

Operations

Publications

Trudeau SJ, Hwang H, Mathur D, Begum K, Petrey D, Murray D, Honig B. <scp>PrePCI</scp> : A structure‐ and chemical similarity‐informed database of predicted protein compound interactions. Protein Science. 2023;32(4). doi:10.1002/pro.4594. PMID:36776141. PMCID:PMC10019447.