Goby framework

Goby framework compresses and organizes high-throughput sequencing (HTS) datasets to reduce storage and computational burdens while preserving data fidelity.


Key Features:

  • Seamless Data Schema Evolution: Supports schema evolution to adapt data representations to new sequencing technologies and analysis methods without extensive reconfiguration.
  • Advanced Compression Techniques: Employs novel compression algorithms that store spliced RNA-Seq alignments at less than 4% of the size of a standard BAM file while maintaining perfect data fidelity and reducing dataset sizes by more than 40% compared to previous compression methods.
  • Multi-Tier Data Organization: Implements a multi-tier organization for collection, analysis, and archiving to minimize storage requirements and reduce local and network computational load.
  • Integration with High-Throughput Sequencing Analyses: Integrates with software suites supporting exome, gene expression, and DNA methylation assays for downstream analysis.

Scientific Applications:

  • HTS data storage and archiving: Reduces storage footprint and facilitates transfer of large-scale sequencing datasets.
  • RNA-Seq alignment preservation: Stores spliced RNA-Seq alignments with perfect fidelity relative to standard BAM representations.
  • Assay-specific analyses: Supports exome, gene expression, and DNA methylation datasets within high-throughput sequencing workflows.
  • Scalable genomics data management: Enables scalable handling of growing HTS data volumes through compression and tiered organization.

Methodology:

Implements schema evolution, novel compression algorithms for HTS data (including compression yielding <4% BAM size for spliced RNA-Seq alignments), and a multi-tier data organization strategy for collection, analysis, and archiving.

Topics

Details

License:
GPL-2.0
Maturity:
Mature
Tool Type:
workflow
Operating Systems:
Linux
Programming Languages:
Java, C++, Python, C
Added:
1/13/2017
Last Updated:
12/10/2018

Operations

Publications

Campagne F, Dorff KC, Chambwe N, Robinson JT, Mesirov JP. Compression of Structured High-Throughput Sequencing Data. PLoS ONE. 2013;8(11):e79871. doi:10.1371/journal.pone.0079871. PMID:24260313. PMCID:PMC3832420.

Documentation