PyBDA
PyBDA enables automated, distributed analysis of large-scale biological datasets for statistical and machine learning investigations in high-performance computing (HPC) environments.
Key Features:
- Automated distributed analysis: Performs automated, distributed analysis of large-scale biological datasets.
- Apache Spark backend: Uses Apache Spark as its backend framework to enable scalable distributed computation.
- Snakemake integration: Integrates Snakemake to schedule and manage computational jobs across distributed HPC clusters.
- HPC support: Tailored for high-performance computing environments to leverage cluster resources.
- Scalability: Scales to handle hundreds of millions of data points and high-dimensional datasets beyond conventional applications.
- Statistical and machine learning methods: Executes common statistical methods and machine learning algorithms on large datasets.
- Proven at scale: Demonstrated on image-based RNA interference (RNAi) data comprising 150 million single cells.
Scientific Applications:
- Image-based RNAi single-cell analysis: Analysis of image-based RNA interference (RNAi) datasets at single-cell scale, demonstrated on 150 million single cells.
- High-dimensional data analysis: Analysis of high-dimensional bioinformatics datasets comprising hundreds of millions of data points.
- Large-scale statistical and machine learning studies: Application of statistical methods and machine learning algorithms to large biological datasets in HPC settings.
Methodology:
Uses Apache Spark as the backend and integrates Snakemake to schedule and manage computational jobs across distributed HPC clusters to perform automated distributed analysis of large-scale datasets.
Topics
Details
- License:
- GPL-3.0
- Maturity:
- Emerging
- Cost:
- Free of charge
- Tool Type:
- command-line tool
- Operating Systems:
- Linux, Mac
- Programming Languages:
- Python
- Added:
- 11/17/2019
- Last Updated:
- 11/17/2019
Operations
Data Inputs & Outputs
Publications
Dirmeier S, Emmenlauer M, Dehio C, Beerenwinkel N. PyBDA: a command line tool for automated analysis of big biological data sets. BMC Bioinformatics. 2019;20(1). doi:10.1186/s12859-019-3087-8. PMID:31718539. PMCID:PMC6849186.
Documentation
Downloads
- Software packagehttps://pybda.readthedocs.io/en/latest/usage.html
- Source codehttps://github.com/cbg-ethz/pybda