ESLpred
ESLpred predicts the subcellular localization of eukaryotic proteins to aid functional genome annotation.
Key Features:
- Machine learning algorithm: Uses Support Vector Machines (SVM) for classification of protein subcellular localization.
- Feature-based prediction: Incorporates amino acid composition, dipeptide composition, and physico-chemical properties such as hydrophobicity and charge into the predictive model.
- Hybrid input vector: Combines features into a 458-dimensional input vector comprising 400 dipeptide composition values, 33 physico-chemical properties, 20 amino acid composition values, and 5 PSI-BLAST output features.
- PSI-BLAST integration: Employs Position-Specific Iterated BLAST (PSI-BLAST) to compare query sequences against a dataset of experimentally annotated proteins and produce output features used in prediction.
- Accuracy metrics: Reports per-location accuracies of 95.3% (nuclear), 85.2% (cytoplasmic), 68.2% (mitochondrial), and 88.9% (extracellular), overall hybrid accuracy of 88.0%, and comparative accuracies of 78.1% (amino acid composition), 77.8% (physico-chemical properties), and 82.9% (dipeptide composition).
- Reliability index: Provides a reliability index (≥3) with 73.5% of predictions achieving an accuracy of 96.4% under that threshold.
- Validation technique: Evaluated using 5-fold cross-validation to assess robustness and generalizability.
Scientific Applications:
- Genomics: Supports functional genome annotation by predicting protein subcellular localization.
- Proteomics: Provides localization information relevant to protein identification and contextual interpretation in proteomic studies.
- Molecular biology: Informs experimental design by indicating probable subcellular compartments for proteins.
- Protein engineering: Guides design considerations by revealing native localization constraints.
- Disease mechanism studies: Aids interpretation of protein mislocalization and its potential role in disease.
- Functional and interaction network analysis: Helps infer protein function and interaction contexts through predicted localization.
Methodology:
ESLpred uses Support Vector Machines trained on features including amino acid composition, dipeptide composition, physico-chemical properties and 5 PSI-BLAST output features combined into a 458-dimensional vector, with performance assessed by 5-fold cross-validation.
Topics
Details
- Tool Type:
- web application
- Operating Systems:
- Linux, Windows, Mac
- Added:
- 2/10/2017
- Last Updated:
- 11/24/2024
Operations
Publications
Bhasin M, Raghava GPS. ESLpred: SVM-based method for subcellular localization of eukaryotic proteins using dipeptide composition and PSI-BLAST. Nucleic Acids Research. 2004;32(Web Server):W414-W419. doi:10.1093/nar/gkh350. PMID:15215421. PMCID:PMC441488.