Mal-Prec
Mal-Prec predicts malonylated sites in protein sequences to identify lysine malonylation relevant to studies of post-translational modification and disease associations such as Type 2 Diabetes Mellitus and cancer.
Key Features:
- Machine Learning Framework: Mal-Prec employs a supervised Support Vector Machine (SVM) classifier to predict malonylation sites.
- One-hot Encoding: Uses one-hot encoding to represent amino acid identities as sequence features.
- Physicochemical Properties (AAindex): Incorporates AAindex-derived physicochemical properties of amino acids as feature inputs.
- Composition of k-spaced Amino Acid Pairs (CKSAAP): Utilizes CKSAAP to capture local sequence context and spacing of residue pairs.
- Feature Selection (PCA): Applies Principal Component Analysis (PCA) to reduce dimensionality and select optimal feature subsets.
- Validation and Performance: Evaluated using five-fold cross-validation and independent test datasets with reported AUCs of 96.47% and 90.72%, respectively, and reported to outperform existing prediction tools.
Scientific Applications:
- Identification of Malonylation Sites: Enables prediction and discovery of novel protein malonylation sites.
- Biomarker and Therapeutic Target Discovery: Supports identification of potential biomarkers and therapeutic targets related to malonylation in diseases such as Type 2 Diabetes Mellitus and cancer.
- Post-translational Modification Research: Facilitates studies in bioinformatics and molecular biology focused on post-translational modifications.
Methodology:
Protein sequences are encoded using one-hot encoding, AAindex physicochemical properties, and CKSAAP; PCA is applied for feature optimization; an SVM classifier is trained and evaluated with five-fold cross-validation and independent test datasets.
Topics
Details
- Tool Type:
- command-line tool
- Programming Languages:
- C, MATLAB
- Added:
- 1/18/2021
- Last Updated:
- 2/19/2021
Operations
Publications
Liu X, Wang L, Li J, Hu J, Zhang X. Mal-Prec: computational prediction of protein Malonylation sites via machine learning based feature integration. BMC Genomics. 2020;21(1). doi:10.1186/s12864-020-07166-w. PMID:33225896. PMCID:PMC7682087.
PMID: 33225896
PMCID: PMC7682087
Funding: - Xuzhou Science and Technology Project: KC17123
- Jiangsu Postdoctoral Science Foundation: 1601080C, 1701062B
- Jiangsu University Natural Science Foundation: 17KJB310015
- Research Foundation for Talented Scholars in Xuzhou Medical University: D2015001