HIVE

HIVE manages and analyzes next-generation sequencing (NGS) data using a distributed storage and compute environment for large-scale genomic data processing.


Key Features:

  • Distributed storage and compute environment: Uses distributed storage and compute resources to handle large NGS datasets.
  • Cloud control server virtualization: Virtualizes services rather than processes, creating an abstraction layer between computational requests and operating system processes.
  • Computation-to-data paradigm: Moves computations to the data instead of transferring data to computational nodes to reduce network and hardware load.
  • Honeycomb data model: Integrates metadata into an object-oriented framework to represent and manage diverse data types.
  • Unified API: Provides a single application programming interface for searching, viewing, and manipulating diverse data types.
  • Hierarchical access control and permissions: Implements fine-grained hierarchical access control to manage data access privileges.
  • NGS dataset operations: Supports deposition, retrieval, annotation, and computational analysis of NGS datasets.
  • Extensible data type integration: Simplifies incorporation of new data types while minimizing database restructuring.

Scientific Applications:

  • NGS data management and analysis: Storage, annotation, and large-scale computational analysis of next-generation sequencing datasets.
  • Scalable high-throughput workflows: Enables massively parallel NGS computations by minimizing data movement and leveraging distributed compute resources.
  • Metadata-driven integration: Facilitates integrated information systems through metadata integration and unified API-based access to diverse data types.
  • Secure controlled-access analyses: Supports fine-grained, hierarchical permissioning suitable for regulated and collaborative analysis environments.

Methodology:

HIVE employs a distributed storage and compute environment with a cloud control server that virtualizes services rather than processes to create an abstraction layer between computational requests and OS processes, moves computations to where data reside to reduce data transfer, uses a honeycomb object-oriented data model integrating metadata, exposes a unified API for data access, and implements a hierarchical access control and permission system.

Topics

Details

Tool Type:
web application
Operating Systems:
Linux, Windows, Mac
Added:
5/19/2018
Last Updated:
12/10/2018

Operations

Publications

Simonyan V, Chumakov K, Dingerdissen H, Faison W, Goldweber S, Golikov A, Gulzar N, Karagiannis K, Vinh Nguyen Lam P, Maudru T, Muravitskaja O, Osipova E, Pan Y, Pschenichnov A, Rostovtsev A, Santana-Quintero L, Smith K, Thompson EE, Tkachenko V, Torcivia-Rodriguez J, Voskanian A, Wan Q, Wang J, Wu T, Wilson C, Mazumder R. High-performance integrated virtual environment (HIVE): a robust infrastructure for next-generation sequence data analysis. Database. 2016;2016:baw022. doi:10.1093/database/baw022. PMID:26989153. PMCID:PMC4795927.

Documentation