StandEnA

StandEnA generates customizable standardized protein annotation databases and presence-absence matrices by retrieving protein sequences and synonyms from multiple public databases based on user-defined standard names to support prokaryotic genome and metagenome comparative analyses.


Key Features:

  • Custom Database Generation: Retrieves protein sequences from multiple public databases based on a user-defined list of standard names and compiles them into a custom database.
  • Synonym Mapping: Searches for and collects synonyms corresponding to standard names across selected public databases to expand annotation coverage.
  • Standardization and Comparability: Produces standardized annotation outputs and reference files containing standard database identifiers to enable comparability across datasets.
  • Presence-Absence Matrix Creation: Generates presence-absence matrices from annotated genomes for comparative analyses of protein presence across samples or conditions.
  • Reference File Generation: Compiles standardized reference files that map standard names to database identifiers for downstream analyses.

Scientific Applications:

  • Prokaryotic Genome Annotation: Supports annotation workflows for prokaryotic genomes by providing custom databases and standardized identifiers.
  • Comparative Genomics and Metagenomics: Enables comparative analyses using presence-absence matrices to identify protein distribution across genomes or metagenome-assembled genomes.
  • Cross-study Standardization: Provides standardized reference files to facilitate consistent data representation and interpretation across studies.
  • Demonstrated Use Case: Applied to six metagenome-assembled genomes to analyze three distinct biological pathways.

Methodology:

Users define a list of standard names; StandEnA retrieves corresponding protein sequences and synonyms from selected public databases and compiles them into a custom database, which is then used for genome annotation and generation of presence-absence matrices and standardized reference files.

Topics

Details

License:
CC-BY-4.0
Cost:
Free of charge
Tool Type:
workflow
Programming Languages:
Python, Perl, Shell
Added:
2/5/2024
Last Updated:
2/5/2024

Operations

Publications

Chafra F, Borim Correa F, Oni F, Konu Karakayalı Ö, Stadler PF, Nunes da Rocha U. StandEnA: a customizable workflow for standardized annotation and generating a presence–absence matrix of proteins. Bioinformatics Advances. 2023;3(1). doi:10.1093/bioadv/vbad069. PMID:37448812. PMCID:PMC10336186.

PMID: 37448812
Funding: - German Research Foundation: 460129525