PoSSuMsearch2

PoSSuMsearch2 accelerates database searches for protein family classification by using position-specific scoring matrices family models (PSSM-FMs) to prefilter sequences prior to profile hidden Markov model (pHMM) searches, reducing computational cost while retaining pHMM sensitivity.


Key Features:

  • Search Space Reduction: Employs PSSM-FMs to prefilter sequences, reducing candidates for pHMM-based searches and retaining over 99.5% of original findings.
  • Speed and Efficiency: Combines full-text indexing, exact p-value computation of PSSM match scores, and fast fragment chaining to achieve more than a 64-fold speedup compared to hmmsearch.
  • High Performance: In experiments, reduced sequences needing pHMM analysis to 0.80% of all sequences, yielding an overall speedup of 43 over unfiltered searches and a lossless filtering speedup of 92 for hmmsearch on UniProtKB/Swiss-Prop.

Scientific Applications:

  • Large-scale genomic analyses: Enables rapid and accurate protein family classification in large-scale genomic datasets.
  • Genome annotation pipelines: Prefilters sequence candidates to reduce the computational cost of pHMM-based genome annotation.

Methodology:

Prefilters sequences using PSSM-FMs; uses full-text indexing, exact p-value computation of PSSM match scores, and fast fragment chaining; then performs subsequent pHMM-based searches (e.g., hmmsearch) on the reduced candidate set.

Topics

Collections

Details

Maturity:
Mature
Cost:
Free of charge (with restrictions)
Tool Type:
command-line tool
Operating Systems:
Linux, Mac
Added:
1/20/2016
Last Updated:
11/25/2024

Operations

Data Inputs & Outputs

Sequence analysis

Publications

Beckstette M, Homann R, Giegerich R, Kurtz S. Significant speedup of database searches with HMMs by search space reduction with PSSM family models. Bioinformatics. 2009;25(24):3251-3258. doi:10.1093/bioinformatics/btp593. PMID:19828575. PMCID:PMC2788931.

Documentation

Downloads

Links