tatajuba
Tatajuba identifies and classifies homopolymeric tracts within sequencing reads to quantify length variation that can drive phase variation and alter bacterial gene expression.
Key Features:
- Automated identification: Detects potential homopolymeric tracts within sequencing data, addressing artefacts from repetitive sequences and decreased base diversity.
- Phenotypic impact analysis: Assesses putative effects of tract length variation on phenotype by linking tract variation to gene expression changes associated with phase variation.
- Polymorphism detection: Highlights polymorphic homopolymeric tracts across samples to reveal genetic variability relevant to bacterial adaptation.
- Scalability: Performs exhaustive exploration of homopolymeric features across large sequencing datasets.
Scientific Applications:
- Analysis in Campylobacter and Bordetella: Applied to sequencing data from Campylobacter jejuni and three Bordetella species to confirm known associations between homopolymer tract variation and phenotypic impact and to identify additional candidate variable tracts.
Methodology:
Implemented in C for efficient processing of large sequencing datasets.
Topics
Details
- License:
- GPL-3.0
- Cost:
- Free of charge
- Tool Type:
- command-line tool
- Operating Systems:
- Linux
- Programming Languages:
- C, Shell, Other
- Added:
- 10/12/2021
- Last Updated:
- 10/12/2021
Operations
Publications
de Oliveira Martins L, Bloomfield S, Stoakes E, Grant A, Page AJ, Mather AE. Tatajuba ― Exploring the distribution of homopolymer tracts. Unknown Journal. 2021. doi:10.1101/2021.06.02.446710.
Documentation
Downloads
- Command-line specificationhttps://github.com/quadram-institute-bioscience/tatajuba