ChemTB
ChemTB predicts active small molecules that inhibit Mycobacterium tuberculosis (Mtb) using machine learning to support discovery of antitubercular drugs.
Key Features:
- Machine learning algorithms: Implements support vector machines, random forest (RF), extreme gradient boosting (XGBoost), and deep neural networks (DNN) for compound activity prediction.
- Molecular representations: Models are trained on various molecular representations and descriptors to characterize compounds.
- Classification task: Classifies compounds as Mtb inhibitors or noninhibitors.
- Ensemble consensus strategies: Integrates outputs from multiple models using two consensus strategies to improve prediction accuracy.
- Stacking ensemble: Uses a stacking approach combining RF, XGBoost, and DNN that yields the most reliable consensus predictions.
- Predictive performance: Reports an AUC of 0.842 on a 10-fold cross-validated training set and an AUC of 0.942 on an external test set.
- Model interpretability: Applies Shapley additive explanations to relate molecular descriptors to bioactivity and explain feature contributions.
Scientific Applications:
- Screening for Mtb inhibitors: Prioritizes candidate molecules for experimental testing as potential Mtb inhibitors.
- Antitubercular drug discovery: Aids identification of novel drug candidates for tuberculosis treatment.
- Drug-resistance targeting: Supports discovery efforts addressing drug-resistant strains, including extensively drug-resistant tuberculosis (XDR-TB).
- Mechanistic interpretation: Elucidates relationships between molecular descriptors and inhibitory activity to inform SAR analysis.
Methodology:
Support vector machines, random forest (RF), XGBoost, and deep neural networks (DNN) are trained on various molecular representations to classify compounds as Mtb inhibitors or noninhibitors; outputs are integrated via two consensus strategies including a stacking ensemble of RF, XGBoost and DNN; performance is evaluated by 10-fold cross-validation (AUC 0.842) and an external test set (AUC 0.942); Shapley additive explanations are used to relate molecular descriptors to bioactivity.
Topics
Details
- Tool Type:
- web application
- Added:
- 6/14/2021
- Last Updated:
- 8/20/2021
Operations
Publications
Ye Q, Chai X, Jiang D, Yang L, Shen C, Zhang X, Li D, Cao D, Hou T. Identification of active molecules against<i>Mycobacterium tuberculosis</i>through machine learning. Briefings in Bioinformatics. 2021;22(5). doi:10.1093/bib/bbab068. PMID:33822874.