ChemTB

ChemTB predicts active small molecules that inhibit Mycobacterium tuberculosis (Mtb) using machine learning to support discovery of antitubercular drugs.


Key Features:

  • Machine learning algorithms: Implements support vector machines, random forest (RF), extreme gradient boosting (XGBoost), and deep neural networks (DNN) for compound activity prediction.
  • Molecular representations: Models are trained on various molecular representations and descriptors to characterize compounds.
  • Classification task: Classifies compounds as Mtb inhibitors or noninhibitors.
  • Ensemble consensus strategies: Integrates outputs from multiple models using two consensus strategies to improve prediction accuracy.
  • Stacking ensemble: Uses a stacking approach combining RF, XGBoost, and DNN that yields the most reliable consensus predictions.
  • Predictive performance: Reports an AUC of 0.842 on a 10-fold cross-validated training set and an AUC of 0.942 on an external test set.
  • Model interpretability: Applies Shapley additive explanations to relate molecular descriptors to bioactivity and explain feature contributions.

Scientific Applications:

  • Screening for Mtb inhibitors: Prioritizes candidate molecules for experimental testing as potential Mtb inhibitors.
  • Antitubercular drug discovery: Aids identification of novel drug candidates for tuberculosis treatment.
  • Drug-resistance targeting: Supports discovery efforts addressing drug-resistant strains, including extensively drug-resistant tuberculosis (XDR-TB).
  • Mechanistic interpretation: Elucidates relationships between molecular descriptors and inhibitory activity to inform SAR analysis.

Methodology:

Support vector machines, random forest (RF), XGBoost, and deep neural networks (DNN) are trained on various molecular representations to classify compounds as Mtb inhibitors or noninhibitors; outputs are integrated via two consensus strategies including a stacking ensemble of RF, XGBoost and DNN; performance is evaluated by 10-fold cross-validation (AUC 0.842) and an external test set (AUC 0.942); Shapley additive explanations are used to relate molecular descriptors to bioactivity.

Topics

Details

Tool Type:
web application
Added:
6/14/2021
Last Updated:
8/20/2021

Operations

Publications

Ye Q, Chai X, Jiang D, Yang L, Shen C, Zhang X, Li D, Cao D, Hou T. Identification of active molecules against<i>Mycobacterium tuberculosis</i>through machine learning. Briefings in Bioinformatics. 2021;22(5). doi:10.1093/bib/bbab068. PMID:33822874.

PMID: 33822874
Funding: - Natural Science Foundation of Zhejiang Province: LZ19H300001 - National Natural Science Foundation of China: 21,575,128, 81,773,632 - Key R&D Program of Zhejiang Province: 2020C03010