K-Pax

K-Pax applies a Bayesian model-based approach to classify aligned protein sequences into functionally divergent subgroups and identify subgroup-specific residues.


Key Features:

  • Automated Group Identification: Infers the number of evolutionary groups directly from aligned sequence data without pre-specification.
  • Model-Based Approach: Uses a Bayesian model to distinguish conserved sequence regions relevant for clustering from noisy regions.
  • Stochastic Optimization Algorithm: Implements a fast stochastic optimization algorithm to maximize posterior probability for clustering.
  • High Specificity and Sensitivity: Demonstrates high specificity and sensitivity in simulated and real-world clustering tasks.

Scientific Applications:

  • Evolutionary Analysis: Identifies evolutionarily related subgroups to elucidate relationships among proteins.
  • Functional Annotation: Highlights residues associated with subgroup-specific functions and residues near active sites to inform annotation.
  • Data-Driven Insights: Provides objective subgrouping from sequence data without arbitrary group specification to support reproducible analyses.

Methodology:

K-Pax employs a Bayesian framework on aligned sequence data, models conserved versus noisy regions, automatically infers subgroup numbers, and uses a fast stochastic optimization algorithm to maximize the estimated posterior probability.

Topics

Details

Tool Type:
command-line tool
Operating Systems:
Windows
Programming Languages:
MATLAB
Added:
12/18/2017
Last Updated:
12/10/2018

Operations

Publications

Marttinen P, Corander J, Törönen P, Holm L. Bayesian search of functionally divergent protein subgroups and their function specific residues. Bioinformatics. 2006;22(20):2466-2474. doi:10.1093/bioinformatics/btl411.

Documentation

Links