FAST: Fast Analysis of Sequences Toolbox

FAST: Fast Analysis of Sequences Toolbox provides utilities for filtering, transforming, annotating, and analyzing biological sequence and alignment data for bioinformatics and molecular evolutionary studies.


Key Features:

  • GNU Textutils-style utilities: Implements utilities such as fasgrep, fascut, and fastr modeled after grep, cut, and tr for sequence-oriented text processing.
  • Workflow encoding: Supports compact combinatorial encoding of data workflows to represent complex processing steps and enable reproducible command composition.
  • Data transformation and analysis: Automates numerical, taxonomic, and text-based sorting, selection, and transformation of sequence records and alignment sites using content, index ranges, descriptive tags, and annotated features.
  • Inline calculated analytics: Provides on-the-fly calculations such as composition and codon usage for sequence records and alignment sites.
  • Molecular evolutionary analysis: Enables automated extraction of sites based on content and annotated features and supports molecular population genetic statistics.
  • File format support: Accepts Multi-FastA (a restriction of BioPerl FastA) as the default exchange format and supports Sanger and Illumina 1.8+ FastQ files.

Scientific Applications:

  • Site extraction and alignment analysis: Precise extraction and analysis of alignment sites based on content and annotated features for evolutionary inference.
  • Molecular population genetics: Calculation and aggregation of molecular population genetic statistics from sequence datasets.
  • Composition and codon-usage analysis: Assessment of nucleotide/amino-acid composition and codon-usage patterns across sequences or alignment sites.
  • Sequence filtering and annotation workflows: Programmatic selection, filtering, and annotation of sequence records using index ranges, tags, and feature annotations.
  • Bioinformatics workflow prototyping: Rapid construction and testing of compact, scriptable processing pipelines for sequence data manipulation and analysis.

Methodology:

Implemented in Perl using BioPerl; supports Multi-FastA (a restriction of the BioPerl FastA format) and Sanger and Illumina 1.8+ FastQ input formats.

Topics

Details

Maturity:
Mature
Tool Type:
workflow
Operating Systems:
Mac
Programming Languages:
Perl
Added:
1/13/2017
Last Updated:
11/25/2024

Operations

Publications

Lawrence TJ, Kauffman KT, Amrine KCH, Carper DL, Lee RS, Becich PJ, Canales CJ, Ardell DH. FAST: FAST Analysis of Sequences Toolbox. Frontiers in Genetics. 2015;6. doi:10.3389/fgene.2015.00172. PMID:26042145. PMCID:PMC4437040.

Documentation