SCI-PHY

SCI-PHY performs phylogenomic classification and protein subfamily identification to improve function prediction and detect novel subfamilies.


Key Features:

  • Subfamily classification: Automatic identification of protein subfamilies using subfamily Hidden Markov Models (HMMs) constructed from multiple sequence alignments.
  • Subfamily HMM construction: Construction of subfamily HMMs to represent conserved sequence patterns within identified subfamilies.
  • Efficient scoring scheme: Computationally efficient scoring based on family and subfamily HMMs for high-throughput sequence classification and discrimination of novel subfamilies.
  • Logistic regression for novelty detection: Application of logistic regression to distinguish sequences representing novel subfamilies from known ones.
  • Information-sharing protocol: Estimation of subfamily HMM parameters using an information-sharing protocol that leverages conservation patterns across related families and subfamilies.
  • High specificity and novel subtype prediction: Subfamily HMMs increase separation between homologous and non-homologous proteins, yielding high specificity in classification and enabling prediction of novel protein subtypes.
  • Integration with PhyloFacts: Use of pre-calculated subfamily predictions and HMMs from the PhyloFacts resource (Berkeley Phylogenomics Group) covering over 40,000 protein families and domains.

Scientific Applications:

  • Functional annotation: Improving accuracy of protein function prediction for genes with limited experimental characterization.
  • Evolutionary studies: Supporting phylogenomic analyses and elucidation of evolutionary relationships at the subfamily level.
  • Proteome characterization: Enabling detection and classification of novel protein subtypes within proteomic datasets.
  • Database search specificity: Enhancing separation of homologous versus non-homologous proteins during sequence database searches.

Methodology:

Uses multiple sequence alignments to construct subfamily HMMs, applies family- and subfamily-based HMM scoring, employs logistic regression for novelty detection, and estimates HMM parameters via an information-sharing protocol, with evaluations measuring separation between homologous and non-homologous proteins.

Topics

Details

Tool Type:
command-line tool
Operating Systems:
Linux
Programming Languages:
C++, Python
Added:
8/3/2017
Last Updated:
11/25/2024

Operations

Publications

Brown DP, Krishnamurthy N, Sjölander K. Automated Protein Subfamily Identification and Classification. PLoS Computational Biology. 2007;3(8):e160. doi:10.1371/journal.pcbi.0030160. PMID:17708678. PMCID:PMC1950344.

Documentation

Links