The Extensive de novo TE Annotator (EDTA)

The Extensive de novo TE Annotator (EDTA) annotates transposable elements (TEs) de novo in eukaryotic genomes to generate high-quality, non-redundant TE libraries for genome annotation and evolutionary analyses.


Key Features:

  • High-Quality De Novo Assembly: Facilitates high-quality de novo assembly of TEs to enable comprehensive annotations in highly repetitive genomic regions.
  • Benchmarking and Performance Evaluation: Built on rigorous benchmarking against a manually curated rice TE library and evaluates TE annotation methods for LTR retrotransposons, TIR transposons, MITEs, and Helitrons.
  • Performance Metrics: Employs sensitivity, specificity, accuracy, precision, false discovery rate (FDR), and F1 score to assess annotation effectiveness.
  • Non-Redundant TE Library Generation: Produces a filtered, non-redundant library containing both structurally intact and fragmented elements for whole-genome TE annotation.
  • Deconvolution of Nested Insertions: Deconvolutes nested TE insertions to improve clarity and accuracy of annotations.
  • Cross-Species Robustness: Demonstrated robustness across plant (e.g., maize) and animal (e.g., Drosophila) genomes.
  • Integration of Multiple Annotation Programs: Selects and integrates the most robust TE annotation programs into a cohesive pipeline that processes raw TE candidates to produce refined annotations.

Scientific Applications:

  • TE diversity and evolutionary dynamics: Enables study of TE diversity and evolutionary dynamics within and across species by providing comprehensive TE annotations.
  • Genome annotation and structural analysis: Supports generation of high-quality TE libraries for whole-genome TE annotation and investigation of TE impacts on genome structure and function.
  • Method and tool benchmarking: Provides a framework to systematically evaluate and compare TE annotation methods using curated reference libraries and quantitative metrics.

Methodology:

EDTA selects the most robust TE annotation programs based on systematic benchmarking against a manually curated rice TE library, integrates those programs into a cohesive pipeline that processes raw TE candidates, deconvolutes nested insertions, evaluates LTR/TIR/MITE/Helitron annotations using sensitivity, specificity, accuracy, precision, FDR, and F1 score, and outputs a filtered, non-redundant TE library containing intact and fragmented elements.

Topics

Details

Tool Type:
command-line tool
Programming Languages:
Perl, Python, Shell
Added:
1/14/2020
Last Updated:
12/25/2020

Operations

Publications

Ou S, Su W, Liao Y, Chougule K, Agda JRA, Hellinga AJ, Lugo CSB, Elliott TA, Ware D, Peterson T, Jiang N, Hirsch CN, Hufford MB. Benchmarking transposable element annotation methods for creation of a streamlined, comprehensive pipeline. Genome Biology. 2019;20(1). doi:10.1186/s13059-019-1905-y. PMID:31843001. PMCID:PMC6913007.

PMID: 31843001
PMCID: PMC6913007
Funding: - Division of Molecular and Cellular Biosciences: IOS-1546727, IOS-1740874, IOS-1744001 - National Institute of Food and Agriculture: IOW05282