Mal-Prec

Mal-Prec predicts malonylated sites in protein sequences to identify lysine malonylation relevant to studies of post-translational modification and disease associations such as Type 2 Diabetes Mellitus and cancer.


Key Features:

  • Machine Learning Framework: Mal-Prec employs a supervised Support Vector Machine (SVM) classifier to predict malonylation sites.
  • One-hot Encoding: Uses one-hot encoding to represent amino acid identities as sequence features.
  • Physicochemical Properties (AAindex): Incorporates AAindex-derived physicochemical properties of amino acids as feature inputs.
  • Composition of k-spaced Amino Acid Pairs (CKSAAP): Utilizes CKSAAP to capture local sequence context and spacing of residue pairs.
  • Feature Selection (PCA): Applies Principal Component Analysis (PCA) to reduce dimensionality and select optimal feature subsets.
  • Validation and Performance: Evaluated using five-fold cross-validation and independent test datasets with reported AUCs of 96.47% and 90.72%, respectively, and reported to outperform existing prediction tools.

Scientific Applications:

  • Identification of Malonylation Sites: Enables prediction and discovery of novel protein malonylation sites.
  • Biomarker and Therapeutic Target Discovery: Supports identification of potential biomarkers and therapeutic targets related to malonylation in diseases such as Type 2 Diabetes Mellitus and cancer.
  • Post-translational Modification Research: Facilitates studies in bioinformatics and molecular biology focused on post-translational modifications.

Methodology:

Protein sequences are encoded using one-hot encoding, AAindex physicochemical properties, and CKSAAP; PCA is applied for feature optimization; an SVM classifier is trained and evaluated with five-fold cross-validation and independent test datasets.

Topics

Details

Tool Type:
command-line tool
Programming Languages:
C, MATLAB
Added:
1/18/2021
Last Updated:
2/19/2021

Operations

Publications

Liu X, Wang L, Li J, Hu J, Zhang X. Mal-Prec: computational prediction of protein Malonylation sites via machine learning based feature integration. BMC Genomics. 2020;21(1). doi:10.1186/s12864-020-07166-w. PMID:33225896. PMCID:PMC7682087.

PMID: 33225896
PMCID: PMC7682087
Funding: - Xuzhou Science and Technology Project: KC17123 - Jiangsu Postdoctoral Science Foundation: 1601080C, 1701062B - Jiangsu University Natural Science Foundation: 17KJB310015 - Research Foundation for Talented Scholars in Xuzhou Medical University: D2015001