PyBDA

PyBDA enables automated, distributed analysis of large-scale biological datasets for statistical and machine learning investigations in high-performance computing (HPC) environments.


Key Features:

  • Automated distributed analysis: Performs automated, distributed analysis of large-scale biological datasets.
  • Apache Spark backend: Uses Apache Spark as its backend framework to enable scalable distributed computation.
  • Snakemake integration: Integrates Snakemake to schedule and manage computational jobs across distributed HPC clusters.
  • HPC support: Tailored for high-performance computing environments to leverage cluster resources.
  • Scalability: Scales to handle hundreds of millions of data points and high-dimensional datasets beyond conventional applications.
  • Statistical and machine learning methods: Executes common statistical methods and machine learning algorithms on large datasets.
  • Proven at scale: Demonstrated on image-based RNA interference (RNAi) data comprising 150 million single cells.

Scientific Applications:

  • Image-based RNAi single-cell analysis: Analysis of image-based RNA interference (RNAi) datasets at single-cell scale, demonstrated on 150 million single cells.
  • High-dimensional data analysis: Analysis of high-dimensional bioinformatics datasets comprising hundreds of millions of data points.
  • Large-scale statistical and machine learning studies: Application of statistical methods and machine learning algorithms to large biological datasets in HPC settings.

Methodology:

Uses Apache Spark as the backend and integrates Snakemake to schedule and manage computational jobs across distributed HPC clusters to perform automated distributed analysis of large-scale datasets.

Topics

Details

License:
GPL-3.0
Maturity:
Emerging
Cost:
Free of charge
Tool Type:
command-line tool
Operating Systems:
Linux, Mac
Programming Languages:
Python
Added:
11/17/2019
Last Updated:
11/17/2019

Operations

Data Inputs & Outputs

Publications

Dirmeier S, Emmenlauer M, Dehio C, Beerenwinkel N. PyBDA: a command line tool for automated analysis of big biological data sets. BMC Bioinformatics. 2019;20(1). doi:10.1186/s12859-019-3087-8. PMID:31718539. PMCID:PMC6849186.

Documentation

Downloads