iHSP-PseRAAAC
iHSP-PseRAAAC classifies heat shock proteins (HSPs) into their functional families using sequence-derived reduced amino acid alphabet combined with pseudo amino acid composition for sequence-based protein family assignment.
Key Features:
- Target families: Classifies HSPs into six families: HSP20 (small heat shock proteins, sHSP), HSP40 (J-class proteins), HSP60 (GroEL/ES), HSP70, HSP90, and HSP100.
- Reduced amino acid alphabet: Uses a reduced representation of amino acids to lower feature dimensionality and improve model robustness.
- Pseudo amino acid composition (PseAAC): Integrates pseudo amino acid composition to encode sequence information for classification.
- Dimensionality and overfitting mitigation: Combines reduced alphabet and PseAAC specifically to address high-dimensionality and overfitting in statistical prediction models.
- Validation performance: Reported overall success rate exceeding 87% on a stringent benchmark dataset.
- Validation dataset constraint: Benchmarking ensured that no two HSP sequences within the same subset shared more than 40% pairwise sequence identity.
- Biological scope: Applicable to HSPs from both prokaryotes and eukaryotes, supporting studies of protein homeostasis.
Scientific Applications:
- HSP family assignment: Sequence-based classification of HSPs into HSP20/sHSP, HSP40/J-class, HSP60/GroEL/ES, HSP70, HSP90, and HSP100 families.
- Protein homeostasis research: Facilitates analyses of molecular chaperone roles in protein folding and aggregation across prokaryotes and eukaryotes.
- Disease-related studies: Supports investigation of HSP dysfunction implicated in Parkinson's disease, Alzheimer's disease, and cardiovascular diseases.
- General protein classification: Applies the reduced amino acid alphabet plus PseAAC approach to broader protein classification problems.
Methodology:
Feature encoding combines a reduced amino acid alphabet with pseudo amino acid composition (PseAAC) to reduce dimensionality and mitigate overfitting; model performance was validated on a benchmark dataset with a 40% pairwise sequence identity cutoff, yielding >87% overall success.
Topics
Collections
Details
- Tool Type:
- web application
- Operating Systems:
- Linux, Windows, Mac
- Added:
- 8/3/2017
- Last Updated:
- 11/25/2024
Operations
Publications
Feng P, Chen W, Lin H, Chou K. iHSP-PseRAAAC: Identifying the heat shock protein families using pseudo reduced amino acid alphabet composition. Analytical Biochemistry. 2013;442(1):118-125. doi:10.1016/j.ab.2013.05.024. PMID:23756733.