TRASH
TRASH annotates tandem repeats and maps their higher-order structure in assembled nucleotide sequences from long-read DNA sequencing assemblies to enable de novo analysis of megabase-scale satellite repeat arrays, including eukaryotic centromeres.
Key Features:
- De novo repeat identification: Operates without prior knowledge of repeat composition or monomer sequences to detect tandem repeats.
- Input format: Analyzes fasta assembly files from assembled genomes.
- Tandem repeat mapping: Pinpoints regions occupied by tandem repeats within nucleotide assemblies.
- Higher-order structure detection: Identifies and maps patterns of higher-order repetition within repeat arrays.
- Satellite and centromere targeting: Detects megabase-scale satellite repeat arrays characteristic of eukaryotic centromeres.
- Long-read assembly compatibility: Designed for use with assemblies enabled by long-read DNA sequencing technologies.
- Scalability: Applicable to large and diverse genomic datasets, including plant and human assemblies.
Scientific Applications:
- Centromere research: Analysis of centromeric satellite arrays, demonstrated on the Col-CEN genome of Arabidopsis thaliana and the human CHM13 genome.
- Satellite repeat characterization: Detailed annotation of megabase-scale satellite repeat arrays and their organizational hierarchy.
- Cross-genome repeat analysis: Comparative analysis of tandem repeat structure across diverse genomic assemblies.
Methodology:
Analyzes fasta assembly files to identify regions occupied by tandem repeats and to map repeats and their higher-order structures, operating de novo without requiring prior knowledge of monomer sequences.
Topics
Details
- License:
- MIT
- Cost:
- Free of charge
- Tool Type:
- command-line tool
- Operating Systems:
- Windows, Linux
- Programming Languages:
- R
- Added:
- 1/2/2024
- Last Updated:
- 11/24/2024
Operations
Publications
Wlodzimierz P, Hong M, Henderson IR. TRASH: Tandem Repeat Annotation and Structural Hierarchy. Bioinformatics. 2023;39(5). doi:10.1093/bioinformatics/btad308. PMID:37162382. PMCID:PMC10199239.
PMID: 37162382
PMCID: PMC10199239
Funding: - Biotechnology and Biological Sciences Research Council: BB/S006842/1, BB/S020012/1, BB/V003984/1