FSA

FSA performs probabilistic multiple sequence alignment of protein, RNA, and DNA sequences using pair hidden Markov models and sequence annealing to produce posterior-probability-based alignments and per-column and per-character accuracy and uncertainty estimates for large datasets.


Key Features:

  • Distance-based alignment: Aligns homologous protein, RNA, and DNA sequences using a distance-based approach suitable for thousands of sequences or megabase-long sequences.
  • Statistical motivation and speed: Uses a statistically motivated framework that balances alignment accuracy with computational speed for large datasets.
  • Pair Hidden Markov Models (HMMs): Employs pair HMMs that approximate an insertion/deletion process on a phylogenetic tree and enable computation of posterior probabilities.
  • Sequence annealing algorithm: Integrates posterior probabilities from pair HMMs via a sequence annealing algorithm to construct coherent multiple alignments.
  • Alignment accuracy and uncertainty estimates: Provides estimates of alignment accuracy and uncertainty for every column and character within an alignment based on posterior probabilities.
  • Unsupervised query-specific learning: Incorporates an unsupervised learning procedure tailored to each query to improve parameter estimation.
  • Centroid alignment approach: Uses a centroid alignment strategy to reduce false-positive alignments in biological data.

Scientific Applications:

  • Large-scale sequence analysis: Suitable for alignment tasks in genomics and proteomics involving thousands of sequences or megabase-long regions.
  • Evolutionary biology and phylogenetics: Supports evolutionary inference by providing probabilistic alignments and uncertainty measures relevant to phylogenetic analyses.
  • Comparative genomics: Enables comparison of homologous regions across genomes with posterior-based accuracy assessment.
  • Functional annotation: Assists functional annotation of genes and proteins by supplying statistically supported multiple sequence alignments and confidence estimates.

Methodology:

FSA uses pair hidden Markov models to approximate indel processes on a phylogenetic tree and compute posterior probabilities, applies a sequence-annealing algorithm to integrate these probabilities into multiple alignments, employs unsupervised query-specific learning for parameter estimation, adopts a centroid alignment strategy to reduce false positives, and outputs per-column and per-character accuracy and uncertainty estimates.

Topics

Details

License:
GPL-3.0
Cost:
Free of charge
Tool Type:
command-line tool
Added:
3/21/2022
Last Updated:
11/24/2024

Operations

Publications

Bradley RK, Roberts A, Smoot M, Juvekar S, Do J, Dewey C, Holmes I, Pachter L. Fast Statistical Alignment. PLoS Computational Biology. 2009;5(5):e1000392. doi:10.1371/journal.pcbi.1000392. PMID:19478997. PMCID:PMC2684580.

Documentation

Downloads

Links