SREPRHot
SREPRHot predicts binding hotspots in protein-RNA interactions by identifying residues whose alanine mutation yields a change in binding free energy (ΔΔG) ≥ 2.0 kcal/mol to support analysis of molecular recognition and therapeutic targeting.
Key Features:
- Definition of Binding Hotspots: Binding hotspots are defined as residues whose mutation to alanine results in a binding free energy change (ΔΔG) ≥ 2.0 kcal/mol.
- Data Balancing with SMOTE: The Synthetic Minority Over-sampling Technique (SMOTE) is used to generate synthetic minority-class samples to address dataset imbalance.
- Residue Interface Propensity: A novel feature that captures the propensity of residues to participate in protein-RNA interfaces.
- Topological Features: Features are derived from node-weighted networks to provide structural and functional insights into residues within protein-RNA complexes.
- Random Grouping Feature Selection Strategy: A random grouping strategy combined with a two-step method is applied to determine an optimal feature subset.
- Ensemble Modeling Approach: A stacking ensemble classifier integrates multiple model predictions to improve accuracy and robustness.
- Performance on Independent Test: Reported performance on an independent testing dataset is Sensitivity (SEN) 0.900, Matthews Correlation Coefficient (MCC) 0.557, and Area Under the Curve (AUC) 0.829.
Scientific Applications:
- Structural biology: Identifying binding hotspots to inform interpretation of protein-RNA complex structures.
- Molecular biology: Highlighting functionally critical residues for studies of molecular recognition and mutational effects.
- Drug discovery: Prioritizing residue targets for therapeutic design that modulates protein-RNA interactions.
- Computational biology: Enabling large-scale in silico analysis of protein-RNA interaction hotspots using machine-learning features and ensemble models.
Methodology:
Datasets are balanced using SMOTE; conventional and novel features including residue interface propensity and topological characteristics from node-weighted networks are extracted; a random grouping feature selection strategy with a two-step method refines the feature set; prediction is performed using a stacking ensemble classifier.
Topics
Details
- License:
- Not licensed
- Cost:
- Free of charge
- Tool Type:
- command-line tool
- Operating Systems:
- Mac, Linux, Windows
- Programming Languages:
- Python
- Added:
- 6/30/2022
- Last Updated:
- 11/24/2024
Operations
Publications
Zhou T, Rong J, Liu Y, Gong W, Li C. An ensemble approach to predict binding hotspots in protein–RNA interactions based on SMOTE data balancing and Random Grouping feature selection strategies. Bioinformatics. 2022;38(9):2452-2458. doi:10.1093/bioinformatics/btac138. PMID:35253843.