ChemAGG

ChemAGG predicts colloidal aggregators in chemical libraries to reduce false positives in high-throughput screening (HTS) and support early drug discovery.


Key Features:

  • Data-driven model: Trained on a dataset comprising 12,119 known aggregators and 24,172 drugs or drug candidates.
  • Ensemble learning and representations: Employs two ensemble learning approaches combined with five distinct types of molecular representations to build classification models.
  • Predictive performance: Reported training accuracy 0.950 (AUC 0.987), test accuracy 0.937 (AUC 0.976), and validation on an external set of 5,681 aggregators.
  • Molecular feature analysis: Identified log D, number of hydroxyl groups, aromatic carbons attached to hydrogen atoms, and sulfur atoms in aromatic heterocycles as features indicative of aggregation potential.
  • Rule generalization: Derived simple rules for aggregator detection and validated them against existing druglikeness and aggregation filtering models.
  • Database screening: Applied to multiple curated chemical databases, predicting nearly 20% of evaluated molecules as aggregators.

Scientific Applications:

  • Virtual screening de-risking: Filtering potential aggregators from lead molecules to reduce false positives in HTS.
  • Prioritization for validation: Prioritizing compounds for experimental de-risking based on predicted aggregation propensity.
  • Library assessment: Assessing aggregation risk across curated chemical databases and compound collections.

Methodology:

Models were developed using a dataset of 12,119 aggregators and 24,172 drugs/candidates, applying two ensemble learning approaches with five molecular representations; performance was evaluated using accuracy and AUC on training and test sets and an external validation set of 5,681 aggregators, followed by molecular feature analysis (log D, hydroxyl count, aromatic CH count, sulfur in aromatic heterocycles), derivation of simple rules, validation against existing druglikeness/aggregation filters, and database screening.

Topics

Details

Added:
11/14/2019
Last Updated:
12/10/2020

Operations

Publications

Yang Z, Yang Z, Dong J, Wang L, Zhang L, Ding J, Ding X, Lu A, Hou T, Cao D. Structural Analysis and Identification of Colloidal Aggregators in Drug Discovery. Journal of Chemical Information and Modeling. 2019;59(9):3714-3726. doi:10.1021/acs.jcim.9b00541. PMID:31430151.

PMID: 31430151
Funding: - Ministry of Science and Technology of the People's Republic of China: 2015CB910700 - Natural Science Foundation of?Hunan Province: 2019JJ51003 - Natural Science Foundation of Zhejiang Province: LZ19H300001 - National Science & Technology Major Project of China ?Key New Drug Creation and Manufacturing Program?: 2018ZX09711002-007