HowDe-SBT
HowDe-SBT extends the Sequence Bloom Tree (SBT) framework to improve indexing and query efficiency for large-scale sequence collections such as RNA-seq datasets.
Key Features:
- Efficient Indexing: Implements a novel partitioning strategy that reduces index construction time and space requirements compared to traditional SBT methods.
- Performance Improvements: On real RNA-seq datasets, constructs indexes in less than 36% of the time required by previous SBT methods while using 39% less space.
- Fast Query Processing: Processes small-batch queries at least five times faster than existing SBT approaches.
- Theoretical Framework: Provides an analytical framework to bound and compare space and query performance relative to other SBT methods.
Scientific Applications:
- Transcriptomics (RNA-seq) indexing: Indexes and queries large transcript collections from RNA-seq datasets to support transcript-level analyses.
- Gene expression profiling: Enables rapid retrieval of transcripts for gene expression profiling.
- Differential expression analysis: Facilitates data access required for differential expression analysis across large sample sets.
- Large-scale genomic database querying: Supports efficient indexing and querying of large biological sequence databases for genomic investigations.
Methodology:
Builds on the SBT framework by implementing a partitioning mechanism for information storage and includes a theoretical framework to analyze and bound space and query performance.
Topics
Details
- License:
- MIT
- Tool Type:
- command-line tool
- Programming Languages:
- C++, Python
- Added:
- 11/14/2019
- Last Updated:
- 11/24/2024
Operations
Publications
Harris RS, Medvedev P. Improved representation of sequence bloom trees. Bioinformatics. 2019;36(3):721-727. doi:10.1093/bioinformatics/btz662. PMID:31504157. PMCID:PMC8215923.
PMID: 31504157
PMCID: PMC8215923
Funding: - NSF: CCF-551439057, DBI-1356529, IIS-1453527
- National Institutes of Health: R01GM130691