MaRe

MaRe integrates Docker containers into Apache Spark to enable MapReduce-style parallel execution of existing bioinformatics tools for scalable, data-intensive life sciences analyses.


Key Features:

  • Integration of Docker containers: Encapsulates existing bioinformatics tools as Docker containers to enable their reuse within MapReduce workflows without reimplementation.
  • Apache Spark interoperability: Executes containerized tasks inside the Apache Spark ecosystem to combine Spark's distributed computing with Docker-packaged applications.
  • Scalable data processing: Provides MapReduce-style parallel computation with support for data locality, ingestion from heterogeneous storage systems, and interactive processing capabilities.
  • Parallelization of serial tools: Enables execution of existing serial bioinformatics tools in a parallelized MapReduce fashion to support high-throughput analyses.
  • Demonstrated scalability: Has been validated through application in two data-intensive life science scenarios to show scalable distributed execution.

Scientific Applications:

  • Data-intensive life sciences analysis: Supports large-scale, distributed analyses in life sciences that require MapReduce-style parallelism.
  • High-throughput parallelizable workflows: Facilitates execution of inherently parallelizable, high-throughput bioinformatics tasks by reusing existing optimized tools as containers.
  • Heterogeneous storage integration: Applies to workflows that ingest data from heterogeneous storage systems while maintaining data locality for efficient processing.

Methodology:

MaRe applies the MapReduce programming model on Apache Spark by running Docker containers that encapsulate existing bioinformatics tools to perform parallel computation while maintaining data locality and supporting heterogeneous storage systems.

Topics

Details

License:
Apache-2.0
Tool Type:
library
Programming Languages:
Scala
Added:
1/18/2021
Last Updated:
2/19/2021

Operations

Publications

Capuccini M, Dahlö M, Toor S, Spjuth O. MaRe: Processing Big Data with application containers on Apache Spark. GigaScience. 2020;9(5). doi:10.1093/gigascience/giaa042. PMID:32369166. PMCID:PMC7199472.

PMID: 32369166
PMCID: PMC7199472
Funding: - Horizon 2020: 654241