iPseU-NCP
iPseU-NCP predicts pseudouridine (Ψ) sites in RNA sequences using Random Forest models trained on nucleotide chemical property (NCP) features for computational identification of RNA Ψ-modifications.
Key Features:
- Random Forest algorithm: Uses Random Forest classifiers for prediction of pseudouridine sites.
- Nucleotide Chemical Property (NCP) encoding: Represents RNA sequences with NCP-derived features for model input.
- Benchmark training dataset: Model development used the benchmark dataset from Chen et al. (2016).
- Independent evaluation datasets: Performance was evaluated on H. sapiens (H_200) and S. cerevisiae (S_200) datasets.
- Comparative methods: Performance was compared against iPseU-CNN, PseUI, and iRNA-PseU.
- MCC improvements: Reported increases in Matthew's correlation coefficient (MCC) for H. sapiens of ≈20.0%, 55.0%, and 109.0%, and for S. cerevisiae of ≈6.5%, 35.0%, and 150.0% versus the compared methods.
- Accuracy improvements: Reported accuracy increases for H. sapiens of ≈7.0%, 13.0%, and 20.0%, and for S. cerevisiae of ≈2.0%, 9.5%, and 25.0% versus the compared methods.
Scientific Applications:
- Pseudouridine site identification across RNA types: Applicable to rRNA, mRNA, tRNA, and nuclear/nucleolar RNA pseudouridylation mapping.
- Support for biomedical research: Provides computational evidence relevant to studies in drug discovery, gene therapy, and investigations of human disease-associated RNA modifications.
Methodology:
Random Forest models were trained on the Chen et al. (2016) benchmark dataset using NCP-encoded features and evaluated on independent H. sapiens (H_200) and S. cerevisiae (S_200) datasets with performance measured by Matthew's correlation coefficient (MCC) and accuracy against iPseU-CNN, PseUI, and iRNA-PseU.
Topics
Details
- Tool Type:
- api
- Added:
- 1/18/2021
- Last Updated:
- 2/11/2021
Operations
Publications
Nguyen-Vo T, Nguyen QH, Do TT, Nguyen T, Rahardja S, Nguyen BP. iPseU-NCP: Identifying RNA pseudouridine sites using random forest and NCP-encoded features. BMC Genomics. 2019;20(S10). doi:10.1186/s12864-019-6357-y. PMID:31888464. PMCID:PMC6936030.