EnZymClass

EnZymClass employs ensemble machine learning on alignment-free sequence descriptors to classify proteins and predict enzyme substrate specificity, including plant acyl-ACP thioesterases into short-, long-, and mixed-chain specificities, from small, noisy, and unbalanced datasets.


Key Features:

  • Alignment-Free Feature Extraction: EnZymClass employs 47 alignment-free feature extraction methods that numerically encode protein sequences for classification.
  • Stacked Ensemble Classification Scheme: The method uses a stacked ensemble approach that integrates multiple machine learning models to improve predictive performance over similarity-based methods.
  • Robustness to Limited and Noisy Data: The approach is designed to maintain high prediction performance on small, noisy, and unbalanced datasets.
  • Application to Acyl-ACP Thioesterases (TEs): EnZymClass classifies plant acyl-ACP thioesterases into short-, long-, and mixed-chain free fatty acid specificities.
  • Predictive Performance: Reported performance includes an average accuracy of 0.8 and precision and recall of 0.87 and 0.89, respectively, particularly for medium-chain TE specificity.
  • Discovery of Novel Enzyme Activities: Application to the ThYme database identified two acyl-ACP TEs, ClFatB3 and CwFatB2, with previously uncharacterized activities, and engineering of ClFatB3 increased C10 titers in E. coli hosts.

Scientific Applications:

  • Protein Function Prediction: Predicts enzyme functional categories from sequence data when sequence similarity does not reliably indicate function.
  • Enzyme Discovery and Engineering: Supports discovery of novel thioesterases and guides engineering efforts to modify product chain-length titers such as C10 in microbial hosts.
  • Renewable Chemical Production: Classifies TE substrate specificity to inform selection of enzymes for production of specific free fatty acid chain lengths for renewable chemicals.
  • Functional Annotation with Limited Data: Facilitates annotation and prioritization of proteins in contexts with limited, noisy, or imbalanced experimental datasets, reducing reliance on extensive screening.

Methodology:

EnZymClass extracts 47 alignment-free numerical descriptors from protein sequences and trains a stacked ensemble classifier on a training set of 115 functionally characterized enzyme sequences to predict functions among a pool of 617 sequences, and it was applied to ThYme database sequences for discovery.

Topics

Details

Programming Languages:
Python
Added:
11/27/2021
Last Updated:
11/27/2021

Operations

Publications

Banerjee D, Jindra MA, Linot AJ, Pfleger BF, Maranas CD. EnZymClass: Substrate specificity prediction tool of plant acyl-ACP thioesterases based on Ensemble Learning. Unknown Journal. 2021. doi:10.1101/2021.07.06.451235.