HowDe-SBT

HowDe-SBT extends the Sequence Bloom Tree (SBT) framework to improve indexing and query efficiency for large-scale sequence collections such as RNA-seq datasets.


Key Features:

  • Efficient Indexing: Implements a novel partitioning strategy that reduces index construction time and space requirements compared to traditional SBT methods.
  • Performance Improvements: On real RNA-seq datasets, constructs indexes in less than 36% of the time required by previous SBT methods while using 39% less space.
  • Fast Query Processing: Processes small-batch queries at least five times faster than existing SBT approaches.
  • Theoretical Framework: Provides an analytical framework to bound and compare space and query performance relative to other SBT methods.

Scientific Applications:

  • Transcriptomics (RNA-seq) indexing: Indexes and queries large transcript collections from RNA-seq datasets to support transcript-level analyses.
  • Gene expression profiling: Enables rapid retrieval of transcripts for gene expression profiling.
  • Differential expression analysis: Facilitates data access required for differential expression analysis across large sample sets.
  • Large-scale genomic database querying: Supports efficient indexing and querying of large biological sequence databases for genomic investigations.

Methodology:

Builds on the SBT framework by implementing a partitioning mechanism for information storage and includes a theoretical framework to analyze and bound space and query performance.

Topics

Details

License:
MIT
Tool Type:
command-line tool
Programming Languages:
C++, Python
Added:
11/14/2019
Last Updated:
11/24/2024

Operations

Publications

Harris RS, Medvedev P. Improved representation of sequence bloom trees. Bioinformatics. 2019;36(3):721-727. doi:10.1093/bioinformatics/btz662. PMID:31504157. PMCID:PMC8215923.

PMID: 31504157
PMCID: PMC8215923
Funding: - NSF: CCF-551439057, DBI-1356529, IIS-1453527 - National Institutes of Health: R01GM130691