HASTE
HASTE partitions streaming scientific data by content and assigns "interestingness" scores to prioritize data objects and optimize allocation of storage, computation, and network bandwidth for data-intensive experiments.
Key Features:
- Content-based partitioning: Partitions data streams based on actual content rather than pre-existing metadata.
- Interestingness functions: Employs configurable "interestingness functions" to evaluate the relevance of each data object.
- Interestingness scoring: Assigns an "interestingness score" to each data object to quantify its significance for downstream decisions.
- Policy-driven prioritization: Applies policies that use interestingness scores to determine processing, storage, and transfer priorities.
- Adaptive data pipelines: Builds smart, adaptive pipelines that can be tailored with custom interestingness functions and policies for specific scientific applications.
- Resource optimization: Optimizes utilization of storage, computation, and network bandwidth across the pipeline.
- Edge and cloud deployment: Supports on-premise container cloud processing and public-cloud edge processing, including real-time control loops.
- High-volume stream handling: Manages high-volume, high-velocity microscopy imaging data streams, including transmission electron microscopy (TEM) and high content screening images.
Scientific Applications:
- High Content Screening Experiment: Processes microscopy images in an on-premise container cloud and prioritizes which images to store and compute based on interestingness scores to optimize storage and computation.
- Transmission Electron Microscopy (TEM) Edge Processing: Performs edge processing of TEM images to enable real-time control loops in a public cloud setting and to manage storage, processing, and transfer for high-volume, high-velocity streams.
Methodology:
HASTE dynamically assesses and prioritizes data objects using configurable interestingness functions and policies, assigns per-object interestingness scores, partitions streams by content, and supports edge processing for real-time control loops.
Topics
Details
- License:
- BSD-3-Clause
- Tool Type:
- web application
- Programming Languages:
- Python
- Added:
- 1/18/2021
- Last Updated:
- 1/30/2021
Operations
Publications
Blamey B, Toor S, Dahlö M, Wieslander H, Harrison PJ, Sintorn I, Sabirsh A, Wählby C, Spjuth O, Hellander A. Rapid development of cloud-native intelligent data pipelines for scientific data streams using the HASTE Toolkit. Unknown Journal. 2020. doi:10.1101/2020.09.13.274779.