TieBrush
TieBrush aggregates redundant information from multiple alignment files into condensed BAM representations to summarize large-scale sequencing datasets (RNA-seq, whole-genome, and exome sequencing) and enable rapid visual and computational inspection of transcriptional landscapes across thousands of samples.
Key Features:
- Aggregation: Aggregates redundant aligned-read information from multiple alignment files into a condensed representation.
- BAM condensation: Produces a condensed BAM file that retains much of the original dataset's information while substantially reducing data volume.
- Sequencing support: Supports RNA-seq, whole-genome, and exome sequencing datasets.
- Scalability: Enables analysis across thousands of samples by reducing storage and processing burden through data condensation.
- Statistics extraction: Facilitates extraction of global and subset-specific statistics from aggregated alignments.
- Compatibility: Generates output compatible with most bioinformatics utilities that accept aligned reads as input.
- Visual and computational inspection: Enables rapid visual and computational inspection of transcriptional landscapes across large cohorts.
Scientific Applications:
- Comparative transcriptional analysis: Compare transcriptional landscapes across large cohorts of RNA-seq, whole-genome, or exome samples.
- Large-scale summarization: Summarize and visualize transcriptional patterns across thousands of samples for cohort-level analysis.
- Downstream analysis preparation: Reduce input data volume for downstream computational analyses that require aligned reads.
Methodology:
Aggregates redundant information from multiple alignment files into condensed BAM representations and enables extraction of global and subset-specific statistics for downstream analyses.
Topics
Details
- License:
- MIT
- Tool Type:
- command-line tool
- Programming Languages:
- C++, Python
- Added:
- 12/13/2021
- Last Updated:
- 12/13/2021
Operations
Publications
Varabyou A, Pertea G, Pockrandt C, Pertea M. TieBrush: an efficient method for aggregating and summarizing mapped reads across large datasets. Bioinformatics. 2021;37(20):3650-3651. doi:10.1093/bioinformatics/btab342. PMID:33964128. PMCID:PMC8545345.