K-Pax
K-Pax applies a Bayesian model-based approach to classify aligned protein sequences into functionally divergent subgroups and identify subgroup-specific residues.
Key Features:
- Automated Group Identification: Infers the number of evolutionary groups directly from aligned sequence data without pre-specification.
- Model-Based Approach: Uses a Bayesian model to distinguish conserved sequence regions relevant for clustering from noisy regions.
- Stochastic Optimization Algorithm: Implements a fast stochastic optimization algorithm to maximize posterior probability for clustering.
- High Specificity and Sensitivity: Demonstrates high specificity and sensitivity in simulated and real-world clustering tasks.
Scientific Applications:
- Evolutionary Analysis: Identifies evolutionarily related subgroups to elucidate relationships among proteins.
- Functional Annotation: Highlights residues associated with subgroup-specific functions and residues near active sites to inform annotation.
- Data-Driven Insights: Provides objective subgrouping from sequence data without arbitrary group specification to support reproducible analyses.
Methodology:
K-Pax employs a Bayesian framework on aligned sequence data, models conserved versus noisy regions, automatically infers subgroup numbers, and uses a fast stochastic optimization algorithm to maximize the estimated posterior probability.
Topics
Details
- Tool Type:
- command-line tool
- Operating Systems:
- Windows
- Programming Languages:
- MATLAB
- Added:
- 12/18/2017
- Last Updated:
- 12/10/2018
Operations
Publications
Marttinen P, Corander J, Törönen P, Holm L. Bayesian search of functionally divergent protein subgroups and their function specific residues. Bioinformatics. 2006;22(20):2466-2474. doi:10.1093/bioinformatics/btl411.
Documentation
Training material
http://www.helsinki.fi/bsg/software/K-PAX/supplement.zip