eiR

eiR accelerates structure similarity searching and clustering for large small-molecule compound datasets to enable scalable identification of drug-like compounds.


Key Features:

  • EI-Search: Embeds chemical compounds into a high-dimensional Euclidean space and performs efficient nearest-neighbor searches using locality sensitive hashing (LSH), achieving 40–200× speedups over sequential search while maintaining comparable recall.
  • EI-Clustering: Combines EI-Search with Jarvis-Patrick clustering to cluster very large compound sets, reducing computation from months to days for datasets exceeding 19 million compounds.
  • Embedding and indexing (EI) techniques: Uses embedding and indexing methods to accelerate both similarity search and clustering on large chemical libraries.
  • Benchmarking on large datasets: Demonstrated robustness and efficiency on three datasets ranging from approximately 260,000 to more than 19 million compounds.

Scientific Applications:

  • Large-scale similarity search: Rapid nearest-neighbor searches across millions of small molecules to support compound similarity analyses.
  • Clustering of compound libraries: High-throughput clustering of extensive chemical collections to organize chemical space for downstream analysis.
  • Drug-like compound identification: Scalable analysis of large compound databases to aid identification of drug-like candidates.
  • Analysis of PubChem-scale datasets: Enables similarity searching and clustering applicable to databases the size of the PubChem Compound database.

Methodology:

Compounds are embedded into a high-dimensional Euclidean space and indexed; locality sensitive hashing (LSH) is used for nearest-neighbor search, and EI-Search outputs are combined with Jarvis-Patrick clustering for large-scale clustering, with evaluation on three datasets (~260k to >19M compounds).

Topics

Collections

Details

License:
Artistic-2.0
Tool Type:
command-line tool, library
Operating Systems:
Linux, Windows, Mac
Programming Languages:
R
Added:
1/17/2017
Last Updated:
12/29/2018

Operations

Publications

Cao Y, Jiang T, Girke T. Accelerated similarity searching and clustering of large compound sets by geometric embedding and locality sensitive hashing. Bioinformatics. 2010;26(7):953-959. doi:10.1093/bioinformatics/btq067. PMID:20179075. PMCID:PMC2844998.

Documentation

Downloads

Links