FSA
FSA performs probabilistic multiple sequence alignment of protein, RNA, and DNA sequences using pair hidden Markov models and sequence annealing to produce posterior-probability-based alignments and per-column and per-character accuracy and uncertainty estimates for large datasets.
Key Features:
- Distance-based alignment: Aligns homologous protein, RNA, and DNA sequences using a distance-based approach suitable for thousands of sequences or megabase-long sequences.
- Statistical motivation and speed: Uses a statistically motivated framework that balances alignment accuracy with computational speed for large datasets.
- Pair Hidden Markov Models (HMMs): Employs pair HMMs that approximate an insertion/deletion process on a phylogenetic tree and enable computation of posterior probabilities.
- Sequence annealing algorithm: Integrates posterior probabilities from pair HMMs via a sequence annealing algorithm to construct coherent multiple alignments.
- Alignment accuracy and uncertainty estimates: Provides estimates of alignment accuracy and uncertainty for every column and character within an alignment based on posterior probabilities.
- Unsupervised query-specific learning: Incorporates an unsupervised learning procedure tailored to each query to improve parameter estimation.
- Centroid alignment approach: Uses a centroid alignment strategy to reduce false-positive alignments in biological data.
Scientific Applications:
- Large-scale sequence analysis: Suitable for alignment tasks in genomics and proteomics involving thousands of sequences or megabase-long regions.
- Evolutionary biology and phylogenetics: Supports evolutionary inference by providing probabilistic alignments and uncertainty measures relevant to phylogenetic analyses.
- Comparative genomics: Enables comparison of homologous regions across genomes with posterior-based accuracy assessment.
- Functional annotation: Assists functional annotation of genes and proteins by supplying statistically supported multiple sequence alignments and confidence estimates.
Methodology:
FSA uses pair hidden Markov models to approximate indel processes on a phylogenetic tree and compute posterior probabilities, applies a sequence-annealing algorithm to integrate these probabilities into multiple alignments, employs unsupervised query-specific learning for parameter estimation, adopts a centroid alignment strategy to reduce false positives, and outputs per-column and per-character accuracy and uncertainty estimates.
Topics
Details
- License:
- GPL-3.0
- Cost:
- Free of charge
- Tool Type:
- command-line tool
- Added:
- 3/21/2022
- Last Updated:
- 11/24/2024
Operations
Publications
Bradley RK, Roberts A, Smoot M, Juvekar S, Do J, Dewey C, Holmes I, Pachter L. Fast Statistical Alignment. PLoS Computational Biology. 2009;5(5):e1000392. doi:10.1371/journal.pcbi.1000392. PMID:19478997. PMCID:PMC2684580.
Documentation
Downloads
- Downloads pagehttps://sourceforge.net/projects/fsa/files/