SPDH

SPDH predicts hot spot residues at protein-DNA binding interfaces from sequence information to identify functionally critical residues when three-dimensional structures are unavailable.


Key Features:

  • Feature Extraction: Uses 133 features derived from physicochemical properties, conservation metrics, predicted solvent accessible surface area, and structural characteristics.
  • Feature Selection: Applies systematic feature selection methods to identify an optimal subset of features that improve model performance.
  • Machine Learning Algorithms: Compares support vector machine (SVM), random forest, logistic regression, and k-nearest neighbor classifiers, with variability of physicochemical property features between wild-type and mutant types highlighted as informative.
  • Performance Metrics: Reports evaluation on an independent test set with area under the curve (AUC) of 0.760 and sensitivity of 0.808.

Scientific Applications:

  • Molecular recognition analysis: Identification of critical hot spot residues to advance understanding of protein-DNA interaction mechanisms.
  • Functional site prediction without structures: Prediction of binding-critical residues for proteins lacking experimentally determined three-dimensional structures.
  • Applications in structural biology and computational genomics: Use of sequence-based hot spot predictions to inform studies in structural biology and computational genomics.

Methodology:

Extract 133 sequence-derived features (physicochemical properties, conservation metrics, predicted solvent accessible surface area, structural characteristics), perform systematic feature selection, compare SVM, random forest, logistic regression, and k-nearest neighbor classifiers using variability of physicochemical property features between wild-type and mutant types, and evaluate on an independent test set (AUC 0.760, sensitivity 0.808).

Topics

Details

Programming Languages:
Python
Added:
1/18/2021
Last Updated:
2/20/2021

Operations

Publications

Yao L, Wang H, Bin Y. Predicting Hot Spot Residues at Protein–DNA Binding Interfaces Based on Sequence Information. Interdisciplinary Sciences: Computational Life Sciences. 2020;13(1):1-11. doi:10.1007/s12539-020-00399-z. PMID:33068261.

PMID: 33068261
Funding: - the National Natural Science Foundation of China: 21601001 - the China Postdoctoral Science Foundation: 2018M630699