newDNA-Prot
newDNA-Prot classifies proteins as DNA-binding or non-DNA-binding using an SVM trained on a comprehensive set of sequence-derived and functional features to support studies of gene regulation.
Key Features:
- Support Vector Machine Classifier: Utilizes a support vector machine (SVM) classifier to distinguish DNA-binding from non-DNA-binding proteins.
- Comprehensive Feature Representation: Employs an extensive feature set grouped into six categories: Primary Sequence Based, Evolutionary Profile Based, Predicted Secondary Structure Based, Predicted Relative Solvent Accessibility Based, Physicochemical Property Based, and Biological Function Based.
- Feature Selection Methods: Applies mRMR (Minimum Redundancy Maximum Relevance), a Wrapper method, and a two-stage feature selection method, with the two-stage method reported to outperform the others in removing irrelevant features and reducing redundancy.
- Statistical Analysis: Performs statistical analysis of selected features, with over 95% of features reported as statistically significant across the six feature groups.
Scientific Applications:
- Gene regulation studies: Predicts DNA-binding proteins to support research into gene regulation mechanisms.
- Genomics and molecular biology prediction and benchmarking: Provides predictive performance that outperforms methods such as iDNA-Prot, DNAbinder, and DNA-Prot, with reported improvement of approximately 10% over DNA-Prot across various evaluation measures.
Methodology:
Uses an SVM classifier trained on features from the six specified groups with feature selection by mRMR, Wrapper, and a two-stage method, followed by statistical analysis of selected features (>95% significant).
Topics
Details
- Tool Type:
- command-line tool
- Operating Systems:
- Windows
- Programming Languages:
- Shell
- Added:
- 8/3/2017
- Last Updated:
- 11/25/2024
Operations
Publications
Zhang Y, Xu J, Zheng W, Zhang C, Qiu X, Chen K, Ruan J. newDNA-Prot: Prediction of DNA-binding proteins by employing support vector machine and a comprehensive sequence representation. Computational Biology and Chemistry. 2014;52:51-59. doi:10.1016/j.compbiolchem.2014.09.002. PMID:25240115.
PMID: 25240115
Funding: - National Natural Science Foundation of China: 11201334
- Science and Technology Commission of Tianjin Municipality: 12JCYBJC31900
- Natural Science Fund of China: 10671100, 31050110432, 68075049
- International Development Research Center, Ottawa, Canada: 104519-010