newDNA-Prot

newDNA-Prot classifies proteins as DNA-binding or non-DNA-binding using an SVM trained on a comprehensive set of sequence-derived and functional features to support studies of gene regulation.


Key Features:

  • Support Vector Machine Classifier: Utilizes a support vector machine (SVM) classifier to distinguish DNA-binding from non-DNA-binding proteins.
  • Comprehensive Feature Representation: Employs an extensive feature set grouped into six categories: Primary Sequence Based, Evolutionary Profile Based, Predicted Secondary Structure Based, Predicted Relative Solvent Accessibility Based, Physicochemical Property Based, and Biological Function Based.
  • Feature Selection Methods: Applies mRMR (Minimum Redundancy Maximum Relevance), a Wrapper method, and a two-stage feature selection method, with the two-stage method reported to outperform the others in removing irrelevant features and reducing redundancy.
  • Statistical Analysis: Performs statistical analysis of selected features, with over 95% of features reported as statistically significant across the six feature groups.

Scientific Applications:

  • Gene regulation studies: Predicts DNA-binding proteins to support research into gene regulation mechanisms.
  • Genomics and molecular biology prediction and benchmarking: Provides predictive performance that outperforms methods such as iDNA-Prot, DNAbinder, and DNA-Prot, with reported improvement of approximately 10% over DNA-Prot across various evaluation measures.

Methodology:

Uses an SVM classifier trained on features from the six specified groups with feature selection by mRMR, Wrapper, and a two-stage method, followed by statistical analysis of selected features (>95% significant).

Topics

Details

Tool Type:
command-line tool
Operating Systems:
Windows
Programming Languages:
Shell
Added:
8/3/2017
Last Updated:
11/25/2024

Operations

Publications

Zhang Y, Xu J, Zheng W, Zhang C, Qiu X, Chen K, Ruan J. newDNA-Prot: Prediction of DNA-binding proteins by employing support vector machine and a comprehensive sequence representation. Computational Biology and Chemistry. 2014;52:51-59. doi:10.1016/j.compbiolchem.2014.09.002. PMID:25240115.

PMID: 25240115
Funding: - National Natural Science Foundation of China: 11201334 - Science and Technology Commission of Tianjin Municipality: 12JCYBJC31900 - Natural Science Fund of China: 10671100, 31050110432, 68075049 - International Development Research Center, Ottawa, Canada: 104519-010

Documentation

Links