PROPSEARCH

PROPSEARCH detects functional and structural homologs among protein sequences using compositional property vectors and alignment-free distance measures, enabling identification of homologs even when sequence identity falls below approximately 25%.


Key Features:

  • Novel similarity metric: Defines protein sequence dissimilarity as a weighted sum of differences in compositional properties including singlet and doublet amino acid composition, molecular weight, and isoelectric point.
  • Physico-chemical property analysis: Uses content of bulky and small residues, average hydrophobicity, and average charge to construct compositional vectors for sequences.
  • Vector-based comparison: Transforms query and database sequences into vectors and assesses similarity by computing Euclidean distance between vectors.
  • Weight optimization: Optimizes weights on compositional features to maximize discrimination between structural families.
  • Family averaging: Handles single sequences and multiple sequences merged into an "average" sequence reflecting the average composition of a protein family.
  • Alignment-free database search: Enables rapid scanning of preprocessed databases without performing sequence alignments.
  • Hypothesis generation: Generates hypotheses about potential structural or functional relationships when alignment-based methods fail to detect similarity.

Scientific Applications:

  • Structural family analysis: Identifies members of structural protein families that exhibit low mutual compositional distances when optimized weights are applied.
  • Functional homology detection: Uncovers potential functional homologs missed by alignment-based approaches, supporting annotation and function prediction of novel proteins.

Methodology:

Compute compositional property vectors (singlet and doublet amino acid composition, molecular weight, isoelectric point, bulky/small residue content, average hydrophobicity, average charge); define dissimilarity as a weighted sum and compute Euclidean distance between vectors; optimize weights to maximize discrimination between structural families; support averaging of multiple sequences into a family composition and alignment-free scanning of preprocessed databases.

Topics

Details

Tool Type:
web application
Operating Systems:
Linux, Windows, Mac
Added:
12/6/2017
Last Updated:
11/25/2024

Operations

Publications

Hobohm U, Sander C. A Sequence Property Approach to Searching Protein Databases. Journal of Molecular Biology. 1995;251(3):390-399. doi:10.1006/jmbi.1995.0442. PMID:7650738.

Documentation