Spark-VS

Spark-VS enables massively parallel structure-based virtual screening (SBVS) with Apache Spark to accelerate docking-based screening of large molecular libraries against target receptors.


Key Features:

  • Massively Parallel Processing: Implements distributed execution of docking-based SBVS tasks using Apache Spark to process large compound libraries in parallel.
  • Scalability and Fault Tolerance: Leverages Spark's MapReduce paradigm to provide transparent scalability and fault tolerance across commodity hardware and cloud resources.
  • Cloud Resource Utilization: Optimized for deployment in public cloud environments to utilize scalable compute and storage for large-scale virtual screening.

Scientific Applications:

  • High-throughput Virtual Screening: Enables large-scale docking-based screening of molecular libraries to identify potential lead compounds in drug discovery research.

Methodology:

Adapts existing docking-based SBVS software to operate within a distributed computing framework using Apache Spark; benchmarked by docking a publicly available target receptor against 2.2 million compounds in a cloud environment, demonstrating a parallel efficiency of 87%.

Topics

Details

License:
Apache-2.0
Tool Type:
command-line tool
Operating Systems:
Linux, Windows, Mac
Programming Languages:
Java
Added:
9/1/2018
Last Updated:
11/25/2024

Operations

Publications

Capuccini M, Ahmed L, Schaal W, Laure E, Spjuth O. Large-scale virtual screening on public cloud resources with Apache Spark. Journal of Cheminformatics. 2017;9(1). doi:10.1186/s13321-017-0204-4. PMID:28316653. PMCID:PMC5339264.

Documentation