PeNGaRoo

PeNGaRoo predicts non-classical secreted proteins in Gram-positive bacteria using a two-layer LightGBM ensemble to support studies of protein secretion mechanisms and bacterial pathogenesis.


Key Features:

  • High-Quality Dataset Construction: Built from a curated dataset of experimentally verified non-classical secreted proteins and used to create benchmark datasets for training and validation.
  • Advanced Feature Engineering: Performs comprehensive feature analysis and engineering to assess individual feature performance and improve predictive accuracy beyond traditional methods.
  • Two-Layer LightGBM Ensemble Model: Implements a two-layer ensemble based on LightGBM that integrates multiple single-feature-based models and optimizes parameters via particle swarm optimization.
  • Predictive Performance: Achieves accuracy of 0.900, F-value of 0.903, Matthew's correlation coefficient of 0.803, and area under the curve of 0.963, surpassing previous state-of-the-art predictors.

Scientific Applications:

  • Discovery of Non-Classical Secreted Proteins: Enables identification of non-classical secreted proteins in Gram-positive bacteria to inform studies of bacterial pathogenesis and therapeutic target discovery.
  • Protein Secretion Mechanism Studies and Biotechnology: Supports exploration of protein secretion mechanisms and applications in biotechnology through high predictive accuracy.

Methodology:

Constructs benchmark datasets from experimentally verified non-classical secreted proteins, performs feature analysis and engineering, builds a two-layer ensemble of single-feature-based LightGBM models, and optimizes ensemble parameters using particle swarm optimization.

Topics

Details

Added:
11/14/2019
Last Updated:
1/9/2021

Operations

Publications

Zhang Y, Yu S, Xie R, Li J, Leier A, Marquez-Lago TT, Akutsu T, Smith AI, Ge Z, Wang J, Lithgow T, Song J. PeNGaRoo, a combined gradient boosting and ensemble learning framework for predicting non-classical secreted proteins. Bioinformatics. 2019;36(3):704-712. doi:10.1093/bioinformatics/btz629. PMID:31393553.

PMID: 31393553
Funding: - National Natural Science Foundation of China: 61862017 - Natural Science Foundation of Guangxi: 2016GXNSFCA380005, 2018GXNSFAA138117 - National Health and Medical Research Council of Australia: 1092262, 1127948, 1144652 - Australian Research Council: DP120104460, LP110200333 - National Institute of Allergy and Infectious Diseases of the National Institutes of Health: R01 AI111965 - Outstanding Degree Thesis Cultivation Project of Guilin University of Electronic Technology: 17YJPYSS14 - Australian Laureate Fellow: FL130100038