SCI-PHY
SCI-PHY performs phylogenomic classification and protein subfamily identification to improve function prediction and detect novel subfamilies.
Key Features:
- Subfamily classification: Automatic identification of protein subfamilies using subfamily Hidden Markov Models (HMMs) constructed from multiple sequence alignments.
- Subfamily HMM construction: Construction of subfamily HMMs to represent conserved sequence patterns within identified subfamilies.
- Efficient scoring scheme: Computationally efficient scoring based on family and subfamily HMMs for high-throughput sequence classification and discrimination of novel subfamilies.
- Logistic regression for novelty detection: Application of logistic regression to distinguish sequences representing novel subfamilies from known ones.
- Information-sharing protocol: Estimation of subfamily HMM parameters using an information-sharing protocol that leverages conservation patterns across related families and subfamilies.
- High specificity and novel subtype prediction: Subfamily HMMs increase separation between homologous and non-homologous proteins, yielding high specificity in classification and enabling prediction of novel protein subtypes.
- Integration with PhyloFacts: Use of pre-calculated subfamily predictions and HMMs from the PhyloFacts resource (Berkeley Phylogenomics Group) covering over 40,000 protein families and domains.
Scientific Applications:
- Functional annotation: Improving accuracy of protein function prediction for genes with limited experimental characterization.
- Evolutionary studies: Supporting phylogenomic analyses and elucidation of evolutionary relationships at the subfamily level.
- Proteome characterization: Enabling detection and classification of novel protein subtypes within proteomic datasets.
- Database search specificity: Enhancing separation of homologous versus non-homologous proteins during sequence database searches.
Methodology:
Uses multiple sequence alignments to construct subfamily HMMs, applies family- and subfamily-based HMM scoring, employs logistic regression for novelty detection, and estimates HMM parameters via an information-sharing protocol, with evaluations measuring separation between homologous and non-homologous proteins.
Topics
Details
- Tool Type:
- command-line tool
- Operating Systems:
- Linux
- Programming Languages:
- C++, Python
- Added:
- 8/3/2017
- Last Updated:
- 11/25/2024
Operations
Publications
Brown DP, Krishnamurthy N, Sjölander K. Automated Protein Subfamily Identification and Classification. PLoS Computational Biology. 2007;3(8):e160. doi:10.1371/journal.pcbi.0030160. PMID:17708678. PMCID:PMC1950344.