iHSP-PseRAAAC

iHSP-PseRAAAC classifies heat shock proteins (HSPs) into their functional families using sequence-derived reduced amino acid alphabet combined with pseudo amino acid composition for sequence-based protein family assignment.


Key Features:

  • Target families: Classifies HSPs into six families: HSP20 (small heat shock proteins, sHSP), HSP40 (J-class proteins), HSP60 (GroEL/ES), HSP70, HSP90, and HSP100.
  • Reduced amino acid alphabet: Uses a reduced representation of amino acids to lower feature dimensionality and improve model robustness.
  • Pseudo amino acid composition (PseAAC): Integrates pseudo amino acid composition to encode sequence information for classification.
  • Dimensionality and overfitting mitigation: Combines reduced alphabet and PseAAC specifically to address high-dimensionality and overfitting in statistical prediction models.
  • Validation performance: Reported overall success rate exceeding 87% on a stringent benchmark dataset.
  • Validation dataset constraint: Benchmarking ensured that no two HSP sequences within the same subset shared more than 40% pairwise sequence identity.
  • Biological scope: Applicable to HSPs from both prokaryotes and eukaryotes, supporting studies of protein homeostasis.

Scientific Applications:

  • HSP family assignment: Sequence-based classification of HSPs into HSP20/sHSP, HSP40/J-class, HSP60/GroEL/ES, HSP70, HSP90, and HSP100 families.
  • Protein homeostasis research: Facilitates analyses of molecular chaperone roles in protein folding and aggregation across prokaryotes and eukaryotes.
  • Disease-related studies: Supports investigation of HSP dysfunction implicated in Parkinson's disease, Alzheimer's disease, and cardiovascular diseases.
  • General protein classification: Applies the reduced amino acid alphabet plus PseAAC approach to broader protein classification problems.

Methodology:

Feature encoding combines a reduced amino acid alphabet with pseudo amino acid composition (PseAAC) to reduce dimensionality and mitigate overfitting; model performance was validated on a benchmark dataset with a 40% pairwise sequence identity cutoff, yielding >87% overall success.

Topics

Collections

Details

Tool Type:
web application
Operating Systems:
Linux, Windows, Mac
Added:
8/3/2017
Last Updated:
11/25/2024

Operations

Publications

Feng P, Chen W, Lin H, Chou K. iHSP-PseRAAAC: Identifying the heat shock protein families using pseudo reduced amino acid alphabet composition. Analytical Biochemistry. 2013;442(1):118-125. doi:10.1016/j.ab.2013.05.024. PMID:23756733.

Documentation

Links