PoSSuMsearch2
PoSSuMsearch2 accelerates database searches for protein family classification by using position-specific scoring matrices family models (PSSM-FMs) to prefilter sequences prior to profile hidden Markov model (pHMM) searches, reducing computational cost while retaining pHMM sensitivity.
Key Features:
- Search Space Reduction: Employs PSSM-FMs to prefilter sequences, reducing candidates for pHMM-based searches and retaining over 99.5% of original findings.
- Speed and Efficiency: Combines full-text indexing, exact p-value computation of PSSM match scores, and fast fragment chaining to achieve more than a 64-fold speedup compared to hmmsearch.
- High Performance: In experiments, reduced sequences needing pHMM analysis to 0.80% of all sequences, yielding an overall speedup of 43 over unfiltered searches and a lossless filtering speedup of 92 for hmmsearch on UniProtKB/Swiss-Prop.
Scientific Applications:
- Large-scale genomic analyses: Enables rapid and accurate protein family classification in large-scale genomic datasets.
- Genome annotation pipelines: Prefilters sequence candidates to reduce the computational cost of pHMM-based genome annotation.
Methodology:
Prefilters sequences using PSSM-FMs; uses full-text indexing, exact p-value computation of PSSM match scores, and fast fragment chaining; then performs subsequent pHMM-based searches (e.g., hmmsearch) on the reduced candidate set.
Topics
Collections
Details
- Maturity:
- Mature
- Cost:
- Free of charge (with restrictions)
- Tool Type:
- command-line tool
- Operating Systems:
- Linux, Mac
- Added:
- 1/20/2016
- Last Updated:
- 11/25/2024
Operations
Data Inputs & Outputs
Sequence analysis
Inputs
Outputs
Publications
Beckstette M, Homann R, Giegerich R, Kurtz S. Significant speedup of database searches with HMMs by search space reduction with PSSM family models. Bioinformatics. 2009;25(24):3251-3258. doi:10.1093/bioinformatics/btp593. PMID:19828575. PMCID:PMC2788931.