Sann

Sann predicts protein solvent accessibility using a nearest neighbor method applied to sequence profiles.


Key Features:

  • Nearest Neighbor Methodology: Employs a nearest neighbor approach that leverages z-score values of distance measures within feature vector space to estimate the relative contributions among k-nearest neighbors.
  • Prediction Types: Produces continuous real-value RSA predictions, two-state discrete predictions classifying residues as exposed or buried using a 25% threshold, and three-state discrete predictions with thresholds at 9% and 36% for buried, partially exposed, and fully exposed states.
  • Database Utilization: Uses a Solvent Accessibility Database containing profiles from 5,717 proteins sourced from the PISCES culling server with sequence identity below 25%.
  • Performance Metrics: Achieves 78.38% accuracy for two-state predictions, 65.1% accuracy for three-state predictions, and a Pearson correlation coefficient of 0.676 between predicted and true RSA values for continuous predictions.
  • Benchmarking: On independent CASP8 targets reports 80.89% two-state accuracy, 67.58% three-state accuracy, a Pearson correlation coefficient of 0.727 for continuous predictions, and a mean absolute error of 0.148 for continuous predictions.
  • Scalability: Demonstrates that prediction accuracy improves as the reference database size increases.

Scientific Applications:

  • Protein structure-function analysis: Supports interpretation of residue exposure to inform structure-function relationship studies.
  • Protein folding modeling: Provides solvent accessibility inputs useful for modeling protein folding pathways and conformational states.
  • Stability analysis: Supplies residue-level accessibility information relevant to assessing protein stability and solvent-exposed regions.
  • Identification of functional sites: Aids in locating potentially functional or interaction-prone surface residues based on exposure predictions.

Methodology:

Sann integrates sequence profile data with a nearest neighbor algorithm that computes z-scores of distance measures within feature vector space to weight k-nearest neighbors and produces continuous real-value and discrete two-state/three-state solvent accessibility predictions using a reference database of 5,717 PISCES-filtered protein profiles.

Topics

Details

Tool Type:
web application
Operating Systems:
Linux, Windows, Mac
Programming Languages:
Python
Added:
8/3/2017
Last Updated:
11/25/2024

Operations

Publications

Joo K, Lee SJ, Lee J. Sann: Solvent accessibility prediction of proteins by nearest neighbor method. Proteins: Structure, Function, and Bioinformatics. 2012;80(7):1791-1797. doi:10.1002/prot.24074. PMID:22434533.

PMID: 22434533
Funding: - Creative Research Initiatives of MEST/KOSEF (Center for in-silico Protein Science: 2009-0063610 - National Research Foundation of Korea (NRF) grant funded by the Korea government: 2009-0090085

Documentation

Links