CPE-SLDI
CPE-SLDI improves prediction of coding potential in RNA sequences by applying machine learning and oversampling to address local data imbalance associated with short Open Reading Frames (sORFs) (ORF length < 303 nucleotides) and distinguish coding RNAs from non-coding RNAs (ncRNAs).
Key Features:
- sORF targeting: Specifically addresses sequences containing short Open Reading Frames (sORFs) with ORF length < 303 nucleotides.
- Local data imbalance correction: Handles local imbalance between coding RNAs with sORFs and ncRNAs.
- Oversampling: Employs oversampling techniques to augment datasets for coding RNAs with sORFs.
- Machine learning prediction: Uses advanced machine learning techniques to construct a model distinguishing coding RNAs and ncRNAs.
- Sequence-derived features: Integrates various sequence-derived features from augmented datasets into the prediction model.
- Validation: Demonstrated improved coding potential prediction performance in comparative studies against existing methods.
Scientific Applications:
- Functional genomics: Improving annotation of RNA coding potential in functional genomics studies.
- sORF research: Characterizing RNAs that contain sORFs (ORF length < 303 nt) to clarify their coding potential.
- Disease mechanism exploration: Investigating RNA roles and disease associations by providing more accurate coding versus non-coding classification.
Methodology:
Uses oversampling to augment coding-RNA sORF datasets and constructs a machine learning prediction model that integrates various sequence-derived features from the augmented datasets, with performance evaluated in comparative studies.
Topics
Details
- Tool Type:
- command-line tool
- Programming Languages:
- Python
- Added:
- 1/18/2021
- Last Updated:
- 2/18/2021
Operations
Publications
Chen X, Liu S, Zhang W. Predicting Coding Potential of RNA Sequences by Solving Local Data Imbalance. IEEE/ACM Transactions on Computational Biology and Bioinformatics. 2022;19(2):1075-1083. doi:10.1109/tcbb.2020.3021800. PMID:32886613.