PseU-ST
PseU-ST predicts pseudouridine (Ψ) sites in RNA sequences to identify modification locations for comparative and species-specific analyses.
Key Features:
- Position-aware sequence descriptors: capture positional nucleotide context associated with pseudouridine sites.
- Position-specific trinucleotide propensity based on single strand (PSTNPss): captures positional trinucleotide context associated with Ψ sites.
- Position-specific of three nucleotides (PS3): represents local nucleotide arrangement patterns that contribute to discrimination of modified versus unmodified sites.
- Six optimal RNA sequence encoding schemes (iLearnPlus): selected via systematic evaluation with the iLearnPlus software across a broad set of encoding strategies.
- Feature selection: chi-square filtering and incremental feature selection are used to optimize features for each encoding scheme.
- Machine learning base classifiers: four machine learning algorithms were selected based on preliminary benchmarking and used as base classifiers.
- Stacking ensemble and species-specific optimization: a stacking framework integrates base classifiers and selects optimal feature and classifier combinations for Homo sapiens, Saccharomyces cerevisiae, and Mus musculus through performance comparisons.
Scientific Applications:
- Comparative annotation of pseudouridylation sites: enables comparison of Ψ site distributions across organisms.
- Species-specific annotation: provides species-level prediction models for Homo sapiens, Saccharomyces cerevisiae, and Mus musculus.
- Analysis of RNA modification patterns and functional roles: supplies candidate Ψ sites for downstream analyses of modification patterns and their potential functions.
- Support for mechanistic studies and bioinformatics workflows: facilitates integration of predicted Ψ sites into mechanistic studies in RNA biology and related computational analyses.
Methodology:
Six optimal RNA sequence encoding schemes were selected via systematic evaluation with iLearnPlus; features for each encoding were optimized using chi-square filtering and incremental feature selection; four machine learning algorithms were chosen as base classifiers and integrated via a stacking framework with species-specific selection of optimal feature and classifier combinations for Homo sapiens, Saccharomyces cerevisiae, and Mus musculus through extensive performance comparisons.
Topics
Details
- License:
- Not licensed
- Cost:
- Free of charge
- Tool Type:
- command-line tool
- Operating Systems:
- Mac, Linux, Windows
- Programming Languages:
- Python
- Added:
- 3/18/2023
- Last Updated:
- 11/24/2024
Operations
Publications
Zhang X, Wang S, Xie L, Zhu Y. PseU-ST: A new stacked ensemble-learning method for identifying RNA pseudouridine sites. Frontiers in Genetics. 2023;14. doi:10.3389/fgene.2023.1121694. PMID:36741328. PMCID:PMC9892456.