eiR
eiR accelerates structure similarity searching and clustering for large small-molecule compound datasets to enable scalable identification of drug-like compounds.
Key Features:
- EI-Search: Embeds chemical compounds into a high-dimensional Euclidean space and performs efficient nearest-neighbor searches using locality sensitive hashing (LSH), achieving 40–200× speedups over sequential search while maintaining comparable recall.
- EI-Clustering: Combines EI-Search with Jarvis-Patrick clustering to cluster very large compound sets, reducing computation from months to days for datasets exceeding 19 million compounds.
- Embedding and indexing (EI) techniques: Uses embedding and indexing methods to accelerate both similarity search and clustering on large chemical libraries.
- Benchmarking on large datasets: Demonstrated robustness and efficiency on three datasets ranging from approximately 260,000 to more than 19 million compounds.
Scientific Applications:
- Large-scale similarity search: Rapid nearest-neighbor searches across millions of small molecules to support compound similarity analyses.
- Clustering of compound libraries: High-throughput clustering of extensive chemical collections to organize chemical space for downstream analysis.
- Drug-like compound identification: Scalable analysis of large compound databases to aid identification of drug-like candidates.
- Analysis of PubChem-scale datasets: Enables similarity searching and clustering applicable to databases the size of the PubChem Compound database.
Methodology:
Compounds are embedded into a high-dimensional Euclidean space and indexed; locality sensitive hashing (LSH) is used for nearest-neighbor search, and EI-Search outputs are combined with Jarvis-Patrick clustering for large-scale clustering, with evaluation on three datasets (~260k to >19M compounds).
Topics
Collections
Details
- License:
- Artistic-2.0
- Tool Type:
- command-line tool, library
- Operating Systems:
- Linux, Windows, Mac
- Programming Languages:
- R
- Added:
- 1/17/2017
- Last Updated:
- 12/29/2018
Operations
Publications
Cao Y, Jiang T, Girke T. Accelerated similarity searching and clustering of large compound sets by geometric embedding and locality sensitive hashing. Bioinformatics. 2010;26(7):953-959. doi:10.1093/bioinformatics/btq067. PMID:20179075. PMCID:PMC2844998.