pSuc-Lys
pSuc-Lys predicts lysine succinylation sites in protein sequences by integrating sequence-coupled information into a generalized pseudo amino acid composition framework and using an ensemble of random forest classifiers to improve site identification.
Key Features:
- Sequence Representation: Incorporates sequence-coupled information into a generalized pseudo amino acid composition framework for feature encoding of protein sequences.
- Prediction Methodology: Uses an ensemble predictor composed of multiple individual random forest classifiers to generate site predictions.
- Training Dataset Balancing: Applies random sampling techniques to balance skewed training datasets and mitigate bias in model training.
- Performance: Ensemble learning and balanced training are reported to enhance prediction accuracy and reliability compared to existing methods.
Scientific Applications:
- Basic Research: Identifies potential lysine succinylation sites to elucidate protein functional roles and involvement in cellular processes.
- Drug Development: Characterizes protein succinylation patterns to inform therapeutic strategies for diseases involving dysregulation of post-translational modifications.
Methodology:
Integrates sequence-coupled information into generalized pseudo amino acid composition, applies random sampling to balance training data, and constructs an ensemble predictor from multiple random forest classifiers.
Topics
Details
- Tool Type:
- web application
- Operating Systems:
- Linux, Windows, Mac
- Added:
- 8/3/2017
- Last Updated:
- 11/25/2024
Operations
Publications
Jia J, Liu Z, Xiao X, Liu B, Chou K. pSuc-Lys: Predict lysine succinylation sites in proteins with PseAAC and ensemble random forest approach. Journal of Theoretical Biology. 2016;394:223-230. doi:10.1016/j.jtbi.2016.01.020. PMID:26807806.