ESLpred

ESLpred predicts the subcellular localization of eukaryotic proteins to aid functional genome annotation.


Key Features:

  • Machine learning algorithm: Uses Support Vector Machines (SVM) for classification of protein subcellular localization.
  • Feature-based prediction: Incorporates amino acid composition, dipeptide composition, and physico-chemical properties such as hydrophobicity and charge into the predictive model.
  • Hybrid input vector: Combines features into a 458-dimensional input vector comprising 400 dipeptide composition values, 33 physico-chemical properties, 20 amino acid composition values, and 5 PSI-BLAST output features.
  • PSI-BLAST integration: Employs Position-Specific Iterated BLAST (PSI-BLAST) to compare query sequences against a dataset of experimentally annotated proteins and produce output features used in prediction.
  • Accuracy metrics: Reports per-location accuracies of 95.3% (nuclear), 85.2% (cytoplasmic), 68.2% (mitochondrial), and 88.9% (extracellular), overall hybrid accuracy of 88.0%, and comparative accuracies of 78.1% (amino acid composition), 77.8% (physico-chemical properties), and 82.9% (dipeptide composition).
  • Reliability index: Provides a reliability index (≥3) with 73.5% of predictions achieving an accuracy of 96.4% under that threshold.
  • Validation technique: Evaluated using 5-fold cross-validation to assess robustness and generalizability.

Scientific Applications:

  • Genomics: Supports functional genome annotation by predicting protein subcellular localization.
  • Proteomics: Provides localization information relevant to protein identification and contextual interpretation in proteomic studies.
  • Molecular biology: Informs experimental design by indicating probable subcellular compartments for proteins.
  • Protein engineering: Guides design considerations by revealing native localization constraints.
  • Disease mechanism studies: Aids interpretation of protein mislocalization and its potential role in disease.
  • Functional and interaction network analysis: Helps infer protein function and interaction contexts through predicted localization.

Methodology:

ESLpred uses Support Vector Machines trained on features including amino acid composition, dipeptide composition, physico-chemical properties and 5 PSI-BLAST output features combined into a 458-dimensional vector, with performance assessed by 5-fold cross-validation.

Topics

Details

Tool Type:
web application
Operating Systems:
Linux, Windows, Mac
Added:
2/10/2017
Last Updated:
11/24/2024

Operations

Publications

Bhasin M, Raghava GPS. ESLpred: SVM-based method for subcellular localization of eukaryotic proteins using dipeptide composition and PSI-BLAST. Nucleic Acids Research. 2004;32(Web Server):W414-W419. doi:10.1093/nar/gkh350. PMID:15215421. PMCID:PMC441488.

Documentation