TRASH

TRASH annotates tandem repeats and maps their higher-order structure in assembled nucleotide sequences from long-read DNA sequencing assemblies to enable de novo analysis of megabase-scale satellite repeat arrays, including eukaryotic centromeres.


Key Features:

  • De novo repeat identification: Operates without prior knowledge of repeat composition or monomer sequences to detect tandem repeats.
  • Input format: Analyzes fasta assembly files from assembled genomes.
  • Tandem repeat mapping: Pinpoints regions occupied by tandem repeats within nucleotide assemblies.
  • Higher-order structure detection: Identifies and maps patterns of higher-order repetition within repeat arrays.
  • Satellite and centromere targeting: Detects megabase-scale satellite repeat arrays characteristic of eukaryotic centromeres.
  • Long-read assembly compatibility: Designed for use with assemblies enabled by long-read DNA sequencing technologies.
  • Scalability: Applicable to large and diverse genomic datasets, including plant and human assemblies.

Scientific Applications:

  • Centromere research: Analysis of centromeric satellite arrays, demonstrated on the Col-CEN genome of Arabidopsis thaliana and the human CHM13 genome.
  • Satellite repeat characterization: Detailed annotation of megabase-scale satellite repeat arrays and their organizational hierarchy.
  • Cross-genome repeat analysis: Comparative analysis of tandem repeat structure across diverse genomic assemblies.

Methodology:

Analyzes fasta assembly files to identify regions occupied by tandem repeats and to map repeats and their higher-order structures, operating de novo without requiring prior knowledge of monomer sequences.

Topics

Details

License:
MIT
Cost:
Free of charge
Tool Type:
command-line tool
Operating Systems:
Windows, Linux
Programming Languages:
R
Added:
1/2/2024
Last Updated:
11/24/2024

Operations

Publications

Wlodzimierz P, Hong M, Henderson IR. TRASH: Tandem Repeat Annotation and Structural Hierarchy. Bioinformatics. 2023;39(5). doi:10.1093/bioinformatics/btad308. PMID:37162382. PMCID:PMC10199239.

PMID: 37162382
Funding: - Biotechnology and Biological Sciences Research Council: BB/S006842/1, BB/S020012/1, BB/V003984/1