SAPPHIRE

SAPPHIRE predicts thermophilic proteins from amino acid sequences using a stacking-based ensemble learning framework to improve identification of thermophilic proteins for protein biochemistry and enzyme development.


Key Features:

  • Stacking-based ensemble: Implements a stacking ensemble that integrates outputs from multiple baseline models to produce a meta-predictor (SAPPHIRE).
  • Feature encodings: Uses twelve distinct feature encodings to represent protein sequences without requiring structural information.
  • Baseline models: Trains 72 baseline models by combining twelve feature encodings with six popular machine learning algorithms.
  • Genetic algorithm selection: Employs a genetic algorithm together with a self-assessment-report approach to mine and select informative predicted probabilities from baseline models.
  • Meta-predictor optimization: Builds and refines the final meta-predictor using an optimal feature set derived from selected predictions.
  • Performance evaluation: Achieves 10-fold cross-validation accuracy of 0.942 and Matthew's correlation coefficient (MCC) of 0.884, with reported improvements over existing methods.

Scientific Applications:

  • Large-scale TPP identification: Enables high-throughput prediction of thermophilic proteins from sequence data for proteome-scale screens.
  • Enzyme development: Supports selection of candidate thermophilic enzymes for biotechnology and industrial applications.
  • Protein thermostability research: Facilitates studies in protein biochemistry focused on thermostability determinants using sequence-based predictions.

Methodology:

Train 72 baseline models by combining twelve feature encodings with six machine learning algorithms; apply a genetic algorithm and a self-assessment-report approach to mine and select informative predicted probabilities; construct and optimize a stacking meta-predictor (SAPPHIRE) using an optimal feature set; evaluate performance with 10-fold cross-validation.

Topics

Details

License:
Not licensed
Cost:
Free of charge
Tool Type:
command-line tool
Operating Systems:
Mac, Linux, Windows
Programming Languages:
Python
Added:
9/3/2022
Last Updated:
11/24/2024

Operations

Publications

Charoenkwan P, Schaduangrat N, Moni MA, Lio’ P, Manavalan B, Shoombuatong W. SAPPHIRE: A stacking-based ensemble learning framework for accurate prediction of thermophilic proteins. Computers in Biology and Medicine. 2022;146:105704. doi:10.1016/j.compbiomed.2022.105704. PMID:35690478.

PMID: 35690478
Funding: - Ministry of Science, ICT and Future Planning: 2021R1A2C1014338