Sann
Sann predicts protein solvent accessibility using a nearest neighbor method applied to sequence profiles.
Key Features:
- Nearest Neighbor Methodology: Employs a nearest neighbor approach that leverages z-score values of distance measures within feature vector space to estimate the relative contributions among k-nearest neighbors.
- Prediction Types: Produces continuous real-value RSA predictions, two-state discrete predictions classifying residues as exposed or buried using a 25% threshold, and three-state discrete predictions with thresholds at 9% and 36% for buried, partially exposed, and fully exposed states.
- Database Utilization: Uses a Solvent Accessibility Database containing profiles from 5,717 proteins sourced from the PISCES culling server with sequence identity below 25%.
- Performance Metrics: Achieves 78.38% accuracy for two-state predictions, 65.1% accuracy for three-state predictions, and a Pearson correlation coefficient of 0.676 between predicted and true RSA values for continuous predictions.
- Benchmarking: On independent CASP8 targets reports 80.89% two-state accuracy, 67.58% three-state accuracy, a Pearson correlation coefficient of 0.727 for continuous predictions, and a mean absolute error of 0.148 for continuous predictions.
- Scalability: Demonstrates that prediction accuracy improves as the reference database size increases.
Scientific Applications:
- Protein structure-function analysis: Supports interpretation of residue exposure to inform structure-function relationship studies.
- Protein folding modeling: Provides solvent accessibility inputs useful for modeling protein folding pathways and conformational states.
- Stability analysis: Supplies residue-level accessibility information relevant to assessing protein stability and solvent-exposed regions.
- Identification of functional sites: Aids in locating potentially functional or interaction-prone surface residues based on exposure predictions.
Methodology:
Sann integrates sequence profile data with a nearest neighbor algorithm that computes z-scores of distance measures within feature vector space to weight k-nearest neighbors and produces continuous real-value and discrete two-state/three-state solvent accessibility predictions using a reference database of 5,717 PISCES-filtered protein profiles.
Topics
Details
- Tool Type:
- web application
- Operating Systems:
- Linux, Windows, Mac
- Programming Languages:
- Python
- Added:
- 8/3/2017
- Last Updated:
- 11/25/2024
Operations
Publications
Joo K, Lee SJ, Lee J. Sann: Solvent accessibility prediction of proteins by nearest neighbor method. Proteins: Structure, Function, and Bioinformatics. 2012;80(7):1791-1797. doi:10.1002/prot.24074. PMID:22434533.
DOI: 10.1002/prot.24074
PMID: 22434533
Funding: - Creative Research Initiatives of MEST/KOSEF (Center for in-silico Protein Science: 2009-0063610
- National Research Foundation of Korea (NRF) grant funded by the Korea government: 2009-0090085