DistributedData.jl
DistributedData.jl enables manipulation and analysis of massive, high-dimensional single-cell and phenotyping datasets in distributed computing environments using the Julia programming language.
Key Features:
- Scalability: Processes datasets containing billions of data points and avoids the need for downsampling to preserve analytical integrity.
- High-performance computing: Optimized for high-performance computational resources to enable rapid processing of extremely large datasets, including analysis within minutes.
- Clustering and dimensionality reduction: Provides fast, scalable implementations of clustering and dimensionality reduction techniques tailored for flow and mass cytometry data.
- Distributed parallelization and horizontal scaling: Leverages distributed computing infrastructures to parallelize data-processing tasks and scale horizontally across compute nodes.
- Accuracy preservation: Maintains result quality comparable to state-of-the-art software while scaling to very large datasets.
Scientific Applications:
- Single-cell cytometry analysis: Analysis of high-dimensional flow and mass cytometry datasets at population scale.
- Large-scale clinical and phenotyping studies: Processing and analysis of high-dimensional data generated in extensive clinical and phenotyping cohorts.
- Mouse phenotyping: Applied to massive mouse phenotyping efforts to handle and analyze large single-cell datasets.
- Omics and systems biology: Supports analyses relevant to genomics, proteomics, and systems biology that require scalable handling of high-dimensional data.
Methodology:
Implemented in Julia and explicitly leverages distributed computing infrastructures to parallelize data processing, supports horizontal scaling, and includes fast, scalable implementations of clustering and dimensionality reduction optimized for high-performance computing resources.
Topics
Collections
Details
- License:
- Apache-2.0
- Maturity:
- Mature
- Cost:
- Free of charge
- Tool Type:
- library
- Operating Systems:
- Mac, Linux
- Programming Languages:
- Julia
- Added:
- 2/22/2021
- Last Updated:
- 11/24/2024
Operations
Publications
Kratochvíl M, Hunewald O, Heirendt L, Verissimo V, Vondrášek J, Satagopam VP, Schneider R, Trefois C, Ollert M. GigaSOM.jl: High-performance clustering and visualization of huge cytometry datasets. GigaScience. 2020;9(11). doi:10.1093/gigascience/giaa127. PMID:33205814. PMCID:PMC7672468.