Primer

Primer predicts protein-binding sites within RNA sequences by computing interaction propensities of nucleotide triplets and classifying sites with a support vector machine to inform studies of RNA–protein interactions.


Key Features:

  • Training Set Construction: Removes data redundancy based on nucleotide features rather than sequence similarity to create larger, more informative training datasets.
  • Interaction Propensity Calculation: Calculates interaction propensity (IP) of nucleotide triplets using a newly defined metric applied to an extensive dataset of protein-RNA complexes.
  • Support Vector Machine Model: Uses a support vector machine (SVM) to predict protein-binding sites within RNA sequences, with model development including 5-fold cross-validation on 812 RNA sequences and testing on an independent dataset of 56 RNA sequences.
  • Performance Metrics: Reports 5-fold cross-validation accuracy 86.4%, F-measure 84.8%, Matthews correlation coefficient 0.66 on 812 RNA sequences, and independent-test accuracy 68.1%, F-measure 71.7%, Matthews correlation coefficient 0.35 on 56 RNA sequences.

Scientific Applications:

  • Identification of protein-binding sites: Identifies protein-binding sites within RNA sequences from sequence data alone.
  • Regulatory and functional analysis: Supports analysis of regulatory mechanisms of gene expression, RNA stability, and functional roles of non-coding RNAs.
  • Experimental and therapeutic support: Provides sequence-based insights into RNA-protein interactions to support experimental design and therapeutic target identification.

Methodology:

Removes redundancy from training datasets based on nucleotide features. Computes interaction propensity for nucleotide triplets using a newly defined metric applied to protein-RNA complex datasets. Trains an SVM on the refined training set and validates it via 5-fold cross-validation on 812 RNA sequences and independent testing on 56 RNA sequences.

Topics

Details

Tool Type:
web application
Operating Systems:
Linux, Windows, Mac
Added:
8/3/2017
Last Updated:
11/25/2024

Operations

Publications

Choi S, Han K. Predicting protein-binding RNA nucleotides using the feature-based removal of data redundancy and the interaction propensity of nucleotide triplets. Computers in Biology and Medicine. 2013;43(11):1687-1697. doi:10.1016/j.compbiomed.2013.08.011. PMID:24209914.

PMID: 24209914
Funding: - Ministry of Education, Science and Technology: 2012011982

Documentation

Links