SPDH
SPDH predicts hot spot residues at protein-DNA binding interfaces from sequence information to identify functionally critical residues when three-dimensional structures are unavailable.
Key Features:
- Feature Extraction: Uses 133 features derived from physicochemical properties, conservation metrics, predicted solvent accessible surface area, and structural characteristics.
- Feature Selection: Applies systematic feature selection methods to identify an optimal subset of features that improve model performance.
- Machine Learning Algorithms: Compares support vector machine (SVM), random forest, logistic regression, and k-nearest neighbor classifiers, with variability of physicochemical property features between wild-type and mutant types highlighted as informative.
- Performance Metrics: Reports evaluation on an independent test set with area under the curve (AUC) of 0.760 and sensitivity of 0.808.
Scientific Applications:
- Molecular recognition analysis: Identification of critical hot spot residues to advance understanding of protein-DNA interaction mechanisms.
- Functional site prediction without structures: Prediction of binding-critical residues for proteins lacking experimentally determined three-dimensional structures.
- Applications in structural biology and computational genomics: Use of sequence-based hot spot predictions to inform studies in structural biology and computational genomics.
Methodology:
Extract 133 sequence-derived features (physicochemical properties, conservation metrics, predicted solvent accessible surface area, structural characteristics), perform systematic feature selection, compare SVM, random forest, logistic regression, and k-nearest neighbor classifiers using variability of physicochemical property features between wild-type and mutant types, and evaluate on an independent test set (AUC 0.760, sensitivity 0.808).
Topics
Details
- Programming Languages:
- Python
- Added:
- 1/18/2021
- Last Updated:
- 2/20/2021
Operations
Publications
Yao L, Wang H, Bin Y. Predicting Hot Spot Residues at Protein–DNA Binding Interfaces Based on Sequence Information. Interdisciplinary Sciences: Computational Life Sciences. 2020;13(1):1-11. doi:10.1007/s12539-020-00399-z. PMID:33068261.