TAGADA

TAGADA performs transcript and gene assembly, deconvolution, and analysis to update reference genome annotations by integrating Illumina NovaSeq short reads and PacBio Iso-Seq long reads from RNA-seq projects.


Key Features:

  • Annotation Enhancement: Generates updated reference annotations that incorporate newly identified transcripts, expanding the catalog of genes and transcripts for a species.
  • Expression Computation: Computes expression values for existing and novel annotations to quantify transcript abundance.
  • Identification of lncRNAs: Identifies long non-coding RNAs (lncRNAs) within the transcriptome data.
  • Quality Control Reporting: Produces detailed quality control reports to assess the reliability and accuracy of analysis results.
  • Integration of Multiple Read Types: Integrates Illumina NovaSeq short reads with PacBio Iso-Seq long reads for comprehensive transcriptome profiling.
  • Reproducibility and Execution Framework: Implemented in Nextflow DSL2 and provided as a containerized environment to enable reproducible execution across computing platforms.

Scientific Applications:

  • Genome annotation updating: Integrates new RNA-seq data into existing genome annotations to expand and correct gene models.
  • Demonstrated project application: Applied to RNA-seq data from the GENE-SWiTCH project and reference genomes of chicken and pig, increasing the number of annotated transcripts.
  • Gene model refinement and transcript diversity: Supports refinement of gene models and exploration of transcript diversity across species.

Methodology:

Accepts a genomic sequence, a reference annotation, and RNA-seq reads (Illumina NovaSeq and PacBio Iso-Seq) as inputs and processes them to enhance gene annotations, compute expression values, identify lncRNAs, and generate quality control reports; implemented in Nextflow DSL2 and executed in a containerized environment.

Topics

Details

License:
Apache-2.0
Cost:
Free of charge
Tool Type:
workflow
Operating Systems:
Mac, Linux, Windows
Programming Languages:
Python
Added:
3/8/2024
Last Updated:
11/24/2024

Operations

Publications

Kurylo C, Guyomar C, Foissac S, Djebali S. TAGADA: a scalable pipeline to improve genome annotations with RNA-seq data. NAR Genomics and Bioinformatics. 2023;5(4). doi:10.1093/nargab/lqad089. PMID:37850035. PMCID:PMC10578202.

PMID: 37850035
Funding: - Horizon 2020 Framework Programme: 817998

Links