repeatexplorer2
RepeatExplorer2: Graph-based clustering pipeline for repetitive DNA analysis
RepeatExplorer2 characterizes repetitive sequences in eukaryotic genomes from next-generation sequencing (NGS) short-read data using a graph-based clustering algorithm to identify and classify repetitive DNA elements.
Key Features:
- Graph-based Clustering: Groups short sequence reads by sequence similarity to identify repetitive DNA elements.
- De Novo Repeat Identification: Detects repetitive sequences without prior genome assembly or reference.
- Automated Transposable Element Annotation: Classifies and annotates transposable elements within repeat clusters.
- Tandem Repeat Identification: Detects and characterizes tandemly repeated DNA sequences, including satellite DNA.
- Comparative Analysis: Enables cross-species comparison of repeat composition and abundance.
- ChIP-seq-based Centromeric Repeat Detection: Integrates ChIP-seq data to identify centromeric repeats.
Scientific Applications:
- Comparative Genomics: Analyzes repeat landscapes across multiple species.
- Cytogenetics: Supports development of satellite DNA probes for chromosomal localization studies.
- Genome Characterization: Profiles repetitive DNA content in species lacking high-quality genome assemblies.
Methodology:
RepeatExplorer2 processes NGS short reads by constructing sequence similarity graphs, clustering reads into repeat families, and performing automated annotation of transposable elements and tandem repeats to generate quantitative and structural profiles of repetitive DNA.
Topics
Details
- Added:
- 11/5/2024
- Last Updated:
- 11/5/2024
Operations
Publications
Novák P, Neumann P, Macas J. Global analysis of repetitive DNA from unassembled sequence reads using RepeatExplorer2. Nature Protocols. 2020;15(11):3745-3776. doi:10.1038/s41596-020-0400-y. PMID:33097925.