MetaGraph

MetaGraph indexes and enables full-text search of petabase-scale nucleotide and protein sequence collections using highly compressed annotated De Bruijn-graph-based indexes for large-scale sequence analysis.


Key Features:

  • Efficient Indexing: Uses annotated De Bruijn graphs and advanced data structures and algorithms to create scalable indexes for large sets of DNA or protein sequences.
  • High Compression Ratios: Produces highly compressed indexes with reported compression up to 5800×, allowing the indexed dataset to fit on a single consumer hard drive.
  • Full-Text Search Capability: Enables full-text search across petabases of DNA sequences from all life clades, including viruses, bacteria, fungi, plants, animals, and humans.
  • Cost-Effective On-Demand Search: Supports on-demand searches at an explicit cost of $0.10 per queried Mbp.

Scientific Applications:

  • Biomedical Research: Facilitates large-scale querying and retrieval of sequence data to support diverse biomedical investigations.
  • Integrative Analyses: Enables mining of existing sequence archives for associations and comparative analyses by leveraging full-text search across massive datasets.

Methodology:

Annotated De Bruijn graphs and advanced data structures and algorithms are used to build highly compressed indexes, and the framework supports on-demand searches at $0.10 per queried Mbp.

Topics

Details

Tool Type:
command-line tool, web application
Programming Languages:
C++, Python
Added:
1/18/2021
Last Updated:
11/21/2021

Operations

Publications

Karasikov M, Mustafa H, Danciu D, Zimmermann M, Barber C, Rätsch G, Kahles A. Indexing All Life’s Known Biological Sequences. Unknown Journal. 2020. doi:10.1101/2020.10.01.322164.

Links

Related Tools

rowdiff
Relation: usedBy