Primer
Primer predicts protein-binding sites within RNA sequences by computing interaction propensities of nucleotide triplets and classifying sites with a support vector machine to inform studies of RNA–protein interactions.
Key Features:
- Training Set Construction: Removes data redundancy based on nucleotide features rather than sequence similarity to create larger, more informative training datasets.
- Interaction Propensity Calculation: Calculates interaction propensity (IP) of nucleotide triplets using a newly defined metric applied to an extensive dataset of protein-RNA complexes.
- Support Vector Machine Model: Uses a support vector machine (SVM) to predict protein-binding sites within RNA sequences, with model development including 5-fold cross-validation on 812 RNA sequences and testing on an independent dataset of 56 RNA sequences.
- Performance Metrics: Reports 5-fold cross-validation accuracy 86.4%, F-measure 84.8%, Matthews correlation coefficient 0.66 on 812 RNA sequences, and independent-test accuracy 68.1%, F-measure 71.7%, Matthews correlation coefficient 0.35 on 56 RNA sequences.
Scientific Applications:
- Identification of protein-binding sites: Identifies protein-binding sites within RNA sequences from sequence data alone.
- Regulatory and functional analysis: Supports analysis of regulatory mechanisms of gene expression, RNA stability, and functional roles of non-coding RNAs.
- Experimental and therapeutic support: Provides sequence-based insights into RNA-protein interactions to support experimental design and therapeutic target identification.
Methodology:
Removes redundancy from training datasets based on nucleotide features. Computes interaction propensity for nucleotide triplets using a newly defined metric applied to protein-RNA complex datasets. Trains an SVM on the refined training set and validates it via 5-fold cross-validation on 812 RNA sequences and independent testing on 56 RNA sequences.
Topics
Details
- Tool Type:
- web application
- Operating Systems:
- Linux, Windows, Mac
- Added:
- 8/3/2017
- Last Updated:
- 11/25/2024
Operations
Publications
Choi S, Han K. Predicting protein-binding RNA nucleotides using the feature-based removal of data redundancy and the interaction propensity of nucleotide triplets. Computers in Biology and Medicine. 2013;43(11):1687-1697. doi:10.1016/j.compbiomed.2013.08.011. PMID:24209914.