IDRBP-PPCT

IDRBP-PPCT predicts nucleic acid-binding proteins (NABPs), including DNA-binding proteins (DBPs), RNA-binding proteins (RBPs), and dual RNA–DNA binding proteins (DRBPs), from protein sequences to support analysis of protein–nucleic acid interactions involved in replication, transcription, and translation.


Key Features:

  • PPCT representation: Position-Specific Scoring Matrix (PSSM) and Position-Specific Frequency Matrix (PSFM) Cross Transformation (PPCT) captures evolutionary information embedded within PSSMs and PSFMs and their correlations.
  • Fixed-dimension feature vectors: PPCT transforms protein sequences into fixed-dimension feature vectors suitable for computational analysis.
  • Two-layer random forest framework: A two-layer predictive framework based on the random forest algorithm integrates PPCT-derived features for classification of DBPs, RBPs, and DRBPs.
  • Validation: Model performance was validated using independent datasets and the tomato genome.

Scientific Applications:

  • Protein class identification: Identification of DBPs, RBPs, and DRBPs from protein sequences.
  • Protein–nucleic acid interaction analysis: Investigation of interactions between proteins and nucleic acids relevant to gene expression processes.
  • Gene expression studies: Analysis of proteins involved in replication, transcription, and translation.
  • Genome-wide annotation: Genome-scale prediction and annotation of nucleic acid-binding proteins, as applied to the tomato genome.

Methodology:

PPCT is applied to PSSMs and PSFMs to produce fixed-dimension feature vectors, which are input to a two-layer random forest classifier; performance was evaluated on independent datasets and the tomato genome.

Topics

Details

Tool Type:
web application
Added:
9/27/2021
Last Updated:
9/27/2021

Operations

Data Inputs & Outputs

DNA-binding protein prediction

Publications

Wang N, Zhang J, Liu B. IDRBP-PPCT: Identifying Nucleic Acid-Binding Proteins Based on Position-Specific Score Matrix and Position-Specific Frequency Matrix Cross Transformation. IEEE/ACM Transactions on Computational Biology and Bioinformatics. 2022;19(4):2284-2293. doi:10.1109/tcbb.2021.3069263. PMID:33780341.

PMID: 33780341
Funding: - National Natural Science Foundation of China: 61732012, 61822306 - National Key R&D Program of China: 2018AAA0100100 - Beijing Natural Science Foundation: JQ19019