NanoTRF

NanoTRF identifies high-copy tandem repeats (satellite DNA) in raw Oxford Nanopore Technologies (ONT) sequencing reads to characterize repeat structures and assemble consensus monomer sequences for genomic and cytogenetic analyses.


Key Features:

  • De novo identification: Performs de novo identification of tandem repeats (TRs) from raw reads without requiring prior read assembly, including support for low-depth (<1×) genome sequencing.
  • ONT single-molecule input: Processes raw Oxford Nanopore Technologies (ONT) sequencing reads at single-molecule resolution.
  • Implementation: Implemented as a Python pipeline for processing raw Nanopore reads.
  • TR characterization and consensus assembly: Identifies TRs, characterizes repeat structures, and assembles consensus monomer sequences.
  • Reporting: Generates HTML reports with figures reporting TR genome abundance, monomer sequence, and monomer length.
  • Transposable element annotation: Annotates transposable elements (TEs) within or near satellite DNA arrays.
  • Validation and downstream utility: Validated by fluorescence in situ hybridization (FISH) for evaluation of TR chromosome organization (clustered or dispersed) and provides sequences useful for genome assembly assistance, gap estimation, chromosome mapping, and development of cytogenetic markers.

Scientific Applications:

  • Plant genomics: Characterizing satellite DNA structure, evolution, and chromosomal behavior in plant genomes.
  • Low-depth sequencing analysis: Enabling identification of high-copy TRs from low-depth (<1×) ONT genome sequencing datasets.
  • Cytogenetics and FISH validation: Supporting evaluation of TR chromosome organization (clustered or dispersed) via sequences validated by fluorescence in situ hybridization (FISH).
  • Downstream genomic analyses: Providing consensus monomer sequences for genome assembly assistance, gap estimation, chromosome mapping, and development of cytogenetic markers.
  • Single-molecule repeat studies: Facilitating single-molecule-level studies of repeat organization and dynamics using ONT reads.

Methodology:

Implemented as a Python pipeline that processes raw Nanopore reads to identify tandem repeats, characterize repeat structures, assemble consensus monomer sequences, annotate transposable elements within or near satellite DNA arrays, and generate HTML reports with figures reporting TR genome abundance, monomer sequence, and monomer length.

Topics

Details

License:
MIT
Cost:
Free of charge
Tool Type:
command-line tool
Operating Systems:
Mac, Linux, Windows
Programming Languages:
Python
Added:
10/11/2022
Last Updated:
11/24/2024

Operations

Publications

Kirov I, Kolganova E, Dudnikov M, Yurkevich OY, Amosova AV, Muravenko OV. A Pipeline NanoTRF as a New Tool for De Novo Satellite DNA Identification in the Raw Nanopore Sequencing Reads of Plant Genomes. Plants. 2022;11(16):2103. doi:10.3390/plants11162103. PMID:36015406. PMCID:PMC9413040.

PMID: 36015406
PMCID: PMC9413040
Funding: - Russian Science Foundation: 22-26-00222