CDSbank

CDSbank manages extraction, selection, renaming, and formatting of protein-coding DNA (CDS) sequences and their translated amino acid sequences to support comparative analyses of sequence, structure, function, and evolution.


Key Features:

  • Integrated Sequence Storage: Stores protein-coding DNA (CDS) sequences and corresponding amino acid sequences for each protein annotated in GenBank.
  • Rich Annotation Data: Retains GenBank feature annotations, flags for incomplete 5' and 3' ends, full taxonomic data, and a heuristic ranking of species scientific interest.
  • Automated Dataset Preparation: Provides automated processes to prepare high-quality sequence datasets with configurable defaults and customization.
  • Taxonomy-Aware Processing: Supports taxonomy-aware extraction and labeling of sequences for organism-specific selection and naming.
  • Handling Incomplete Sequences: Identifies and flags incomplete 5' and 3' ends to account for sequence incompleteness in downstream analyses.
  • Cross-Species Sequence Extraction: Extracts synonymous CDS or identical protein sequences across different species based on a single input protein sequence.
  • Property File Creation: Generates labeled property files for annotating or relabeling phylogenetic trees.

Scientific Applications:

  • Comparative Genomics: Assembling homologous CDS and protein sets across taxa for comparative analyses of sequence and function.
  • Molecular Evolution: Selecting taxonomically informed sequence collections and completeness-filtered datasets for evolutionary inference.
  • Functional Annotation: Supporting functional inference by providing CDS and protein sequences with GenBank feature annotations and completeness flags.
  • Phylogenetic Annotation: Producing labeled property files to annotate or relabel phylogenetic trees for interpretation of evolutionary relationships.

Methodology:

Stores CDS and translated amino acid sequences for proteins annotated in GenBank, records GenBank feature annotations and flags for incomplete 5' and 3' ends, includes full taxonomic data and a heuristic species-interest ranking, performs taxonomy-aware extraction and labeling, extracts synonymous CDS or identical proteins across species from a single input sequence, generates labeled property files, and provides automated dataset preparation with configurable defaults.

Topics

Details

License:
GPL-3.0
Tool Type:
web application
Operating Systems:
Linux, Windows, Mac
Programming Languages:
Python
Added:
5/3/2018
Last Updated:
12/10/2018

Operations

Publications

Hazes B. CDSbank: taxonomy-aware extraction, selection, renaming and formatting of protein-coding DNA or amino acid sequences. BMC Bioinformatics. 2014;15(1). doi:10.1186/1471-2105-15-61. PMID:24580755. PMCID:PMC3942066.

Documentation