ChemFLuc

ChemFLuc predicts firefly luciferase (FLuc) inhibitors from chemical structures to identify compounds likely to generate false-positive signals in high-throughput screening (HTS) assays.


Key Features:

  • Extensive Dataset Utilization: Model development used a labeled dataset of 20,888 known FLuc inhibitors and 198,608 non-inhibitors.
  • Machine Learning Algorithms: Employs a combination of three machine learning algorithms, with the best-performing model based on XGBoost.
  • Molecular Descriptors: The best-performing model uses ECFP4 and MOE2d molecular descriptors.
  • Predictive Performance: Achieved balanced accuracy (BA) and area under the ROC curve (AUC) of 0.878 and 0.958 on the validation set, and BA 0.886 and AUC 0.947 on the test set.
  • External Validation: Performance was confirmed on three external validation sets with BA values of 0.864, 0.845, and 0.791.
  • Feature Recognition and Rule Development: Applied Shapley additive explanations (SHAP) to identify structural fragments associated with FLuc inhibition and derived 16 predictive rules with a 70% correction rate.
  • Comparison with Existing Models: Compared models and rules with existing prediction tools in virtual screening contexts and reported superior reliability.
  • Risk Assessment in Chemical Databases: Applied the model to three curated chemical databases and identified approximately 10% of evaluated molecules as potential FLuc inhibitors.

Scientific Applications:

  • HTS Prescreening: Prescreen chemical libraries to flag compounds likely to inhibit firefly luciferase and cause false positives in HTS assays.
  • Virtual Screening Benchmarking: Serve as a comparative model in virtual screening workflows to assess interference from FLuc inhibitors.
  • Chemical Database Risk Assessment: Estimate the prevalence of potential FLuc inhibitors within curated chemical databases to prioritize follow-up validation.

Methodology:

Modeling used a labeled dataset of 20,888 inhibitors and 198,608 non-inhibitors, trained with three machine learning algorithms (best model: XGBoost using ECFP4 and MOE2d descriptors), evaluated by BA and AUC on validation/test sets and three external validation sets, interpreted with SHAP to identify fragments and derive 16 predictive rules, and applied to three curated chemical databases.

Topics

Details

Tool Type:
web application
Added:
1/18/2021
Last Updated:
2/11/2021

Operations

Publications

Yang Z, Dong J, Yang Z, Lu A, Hou T, Cao D. Structural Analysis and Identification of False Positive Hits in Luciferase-Based Assays. Journal of Chemical Information and Modeling. 2020;60(4):2031-2043. doi:10.1021/acs.jcim.9b01188. PMID:32202787.

PMID: 32202787
Funding: - Ministry of Science and Technology of the People's Republic of China: 2015CB910700 - Hong Kong Baptist University: SDF19-0402-P02 - National Natural Science Foundation of China: 21575128, 81773632 - Natural Science Foundation of Zhejiang Province: LZ19H300001 - Zhejiang Province: 2019C03G2010942