PseU-ST

PseU-ST predicts pseudouridine (Ψ) sites in RNA sequences to identify modification locations for comparative and species-specific analyses.


Key Features:

  • Position-aware sequence descriptors: capture positional nucleotide context associated with pseudouridine sites.
  • Position-specific trinucleotide propensity based on single strand (PSTNPss): captures positional trinucleotide context associated with Ψ sites.
  • Position-specific of three nucleotides (PS3): represents local nucleotide arrangement patterns that contribute to discrimination of modified versus unmodified sites.
  • Six optimal RNA sequence encoding schemes (iLearnPlus): selected via systematic evaluation with the iLearnPlus software across a broad set of encoding strategies.
  • Feature selection: chi-square filtering and incremental feature selection are used to optimize features for each encoding scheme.
  • Machine learning base classifiers: four machine learning algorithms were selected based on preliminary benchmarking and used as base classifiers.
  • Stacking ensemble and species-specific optimization: a stacking framework integrates base classifiers and selects optimal feature and classifier combinations for Homo sapiens, Saccharomyces cerevisiae, and Mus musculus through performance comparisons.

Scientific Applications:

  • Comparative annotation of pseudouridylation sites: enables comparison of Ψ site distributions across organisms.
  • Species-specific annotation: provides species-level prediction models for Homo sapiens, Saccharomyces cerevisiae, and Mus musculus.
  • Analysis of RNA modification patterns and functional roles: supplies candidate Ψ sites for downstream analyses of modification patterns and their potential functions.
  • Support for mechanistic studies and bioinformatics workflows: facilitates integration of predicted Ψ sites into mechanistic studies in RNA biology and related computational analyses.

Methodology:

Six optimal RNA sequence encoding schemes were selected via systematic evaluation with iLearnPlus; features for each encoding were optimized using chi-square filtering and incremental feature selection; four machine learning algorithms were chosen as base classifiers and integrated via a stacking framework with species-specific selection of optimal feature and classifier combinations for Homo sapiens, Saccharomyces cerevisiae, and Mus musculus through extensive performance comparisons.

Topics

Details

License:
Not licensed
Cost:
Free of charge
Tool Type:
command-line tool
Operating Systems:
Mac, Linux, Windows
Programming Languages:
Python
Added:
3/18/2023
Last Updated:
11/24/2024

Operations

Publications

Zhang X, Wang S, Xie L, Zhu Y. PseU-ST: A new stacked ensemble-learning method for identifying RNA pseudouridine sites. Frontiers in Genetics. 2023;14. doi:10.3389/fgene.2023.1121694. PMID:36741328. PMCID:PMC9892456.