PreHots
PreHots predicts protein–DNA binding energy hot spots to identify residues that significantly contribute to binding free energy at protein–DNA interfaces.
Key Features:
- Ensemble stacking classifier: PreHots uses an ensemble stacking classifier that integrates multiple machine learning classifiers to predict hot spots.
- Feature selection: The model employs 19 features selected by a sequential backward feature selection algorithm.
- Dataset construction: Two manually curated datasets containing 123 hot spots and 137 non-hot spots from 89 protein–DNA complexes were compiled from literature and databases with redundancy removal for training and validation.
- Performance metrics: In 10-fold cross-validation on the benchmark dataset it achieved sensitivity 0.813 and AUC 0.868, and on an independent test set it achieved sensitivity 0.818 and AUC 0.820.
Scientific Applications:
- Protein–DNA interaction analysis: Predicting hot spots to elucidate residues that dominate binding free energy at protein–DNA interfaces.
- Drug design: Informing the selection of interface residues relevant to targeting protein–DNA interactions.
- Genetic regulation studies: Identifying residues that may modulate regulatory protein–DNA binding.
- Molecular biology research: Supporting studies of molecular mechanisms underlying protein–DNA recognition and binding energetics.
Methodology:
An ensemble stacking classifier integrating multiple machine learning classifiers was trained using 19 features selected by sequential backward feature selection on two manually curated datasets (123 hot spots, 137 non-hot spots from 89 complexes), with performance evaluated by 10-fold cross-validation and an independent test set.
Topics
Details
- Added:
- 1/18/2021
- Last Updated:
- 3/20/2021
Operations
Publications
Pan Y, Zhou S, Guan J. Computationally identifying hot spots in protein-DNA binding interfaces using an ensemble approach. BMC Bioinformatics. 2020;21(S13). doi:10.1186/s12859-020-03675-3. PMID:32938375. PMCID:PMC7495898.