AnnotaPipeline

AnnotaPipeline annotates and validates eukaryotic proteins by integrating transcriptomic (RNA-seq) and proteomic (MS/MS) data with genomic sequences to improve functional annotation accuracy.


Key Features:

  • Integration of Multi-Omics Data: Integrates nucleotide (FASTA), protein, GFF3 structural annotations, RNA-seq (FASTQ), and MS/MS (mzXML) proteomic data to corroborate annotations across multiple evidence types.
  • Proteogenomic Approach: Employs a proteogenomic strategy that combines predicted genomic features with transcriptomic and proteomic evidence to provide transcriptional and expression support for annotations.
  • Automated Workflow: Implements an automated Unix-based pipeline written in Python to process large datasets.
  • Validation of Gene Predictions: Validates predicted genomic features by cross-referencing in silico predictions with experimental transcriptomic and proteomic data to increase annotation confidence.
  • Application to Multiple Organisms: Applied to reannotate genomes of Arabidopsis thaliana, Caenorhabditis elegans, Candida albicans, Trypanosoma cruzi, and Trypanosoma rangeli, increasing the proportion of annotated proteins and reducing hypothetical protein entries compared to public annotations.

Scientific Applications:

  • Functional Genomics: Improving assignment of gene functions across eukaryotic organisms using integrated multi-omics evidence.
  • Comparative Genomics: Providing consistent annotations across species to support comparative genomic analyses.
  • Proteomics Research: Supporting proteomic studies through validated protein identifications and expression evidence from MS/MS data.

Methodology:

Unix-based pipeline implemented in Python that accepts FASTA, GFF3, FASTQ, and mzXML inputs and applies a proteogenomic strategy integrating transcriptomic and proteomic data with computational predictions for experimental validation of genomic features.

Topics

Details

License:
Apache-2.0
Tool Type:
command-line tool, workflow
Operating Systems:
Mac, Linux
Programming Languages:
Python
Added:
2/9/2023
Last Updated:
2/9/2023

Operations

Publications

Maia GA, Filho VB, Kawagoe EK, Teixeira Soratto TA, Moreira RS, Grisard EC, Wagner G. AnnotaPipeline: An integrated tool to annotate eukaryotic proteins using multi-omics data. Frontiers in Genetics. 2022;13. doi:10.3389/fgene.2022.1020100. PMID:36482896. PMCID:PMC9723129.