Berkeley Phylogenomics Group
Berkeley Phylogenomics Group provides phylogenomic algorithms and resources for classifying sequences, constructing multiple sequence alignments and phylogenetic trees, identifying protein subfamilies, and improving protein function prediction.
Key Features:
- PhyloFacts Phylogenomic Encyclopedia: Classifies sequences into pre-computed families and subfamilies using phylogenomic data to improve annotation beyond direct sequence similarity.
- FlowerPower Clustering: Clusters proteins based on shared domain architectures to reveal evolutionary relationships and functional similarities among complex protein families.
- MUSCLE (Multiple Sequence Comparison by Log-Expectation): Produces multiple sequence alignments using kmer counting for fast distance estimation, progressive alignment with a log-expectation score profile function, and refinement via tree-dependent restricted partitioning.
- SATCHMO (Simultaneous Alignment and Tree Construction): Simultaneously constructs phylogenetic trees and multiple sequence alignments, predicts alignable positions across subgroups, and generates profile hidden Markov models at internal nodes to guide alignment and branching order.
- SCI-PHY (Subfamily Identification): Identifies protein subfamilies using relative entropy as a distance metric combined with Dirichlet mixture priors to estimate phylogenetic trees and preserve structurally or functionally important positions.
- PhyloBuilder Web Pipeline: Integrates homolog identification, multiple alignment, phylogenetic tree construction, subfamily identification, and structure prediction starting from a protein sequence.
Scientific Applications:
- Sequence classification and annotation: Assigns proteins to families and subfamilies for more accurate annotation transfer than direct sequence similarity alone.
- Protein function prediction: Improves functional inference by combining phylogenomic classification, subfamily assignment, and structure prediction.
- Subfamily discovery and evolutionary analysis: Detects subfamilies within superfamilies and has been applied to SH2-domain containing proteins to refine subfamily assignments and evolutionary linkages.
- Alignment and tree construction for divergent sequences: Aligns and infers phylogeny for proteins with low sequence identity or structural divergence using simultaneous alignment-tree methods and profile HMMs.
- Domain-architecture-based clustering: Reveals functional and evolutionary relationships among proteins with complex domain arrangements.
Methodology:
Methods explicitly include classification into pre-computed families/subfamilies using phylogenomic data; clustering by shared domain architectures; MUSCLE methods (kmer counting distance estimation, progressive alignment with a log-expectation profile function, tree-dependent restricted partitioning refinement); SATCHMO simultaneous alignment and tree construction with internal-node profile HMMs and alignable-position prediction; SCI-PHY using relative entropy and Dirichlet mixture priors for tree estimation; and PhyloBuilder steps of homolog identification, multiple alignment, phylogenetic tree construction, subfamily identification, and structure prediction.
Topics
Details
- Tool Type:
- web application
- Operating Systems:
- Linux, Windows, Mac
- Added:
- 3/25/2017
- Last Updated:
- 11/25/2024
Operations
Publications
Edgar RC, Sjölander K. SATCHMO: sequence alignment and tree construction using hidden Markov models. Bioinformatics. 2003;19(11):1404-1411. doi:10.1093/bioinformatics/btg158. PMID:12874053.
Edgar RC. MUSCLE: multiple sequence alignment with high accuracy and high throughput. Nucleic Acids Research. 2004;32(5):1792-1797. doi:10.1093/nar/gkh340. PMID:15034147. PMCID:PMC390337.
Glanville JG, Kirshner D, Krishnamurthy N, Sjolander K. Berkeley Phylogenomics Group web servers: resources for structural phylogenomic analysis. Nucleic Acids Research. 2007;35(Web Server):W27-W32. doi:10.1093/nar/gkm325. PMID:17488835. PMCID:PMC1933202.
Sjölander K. Phylogenetic inference in protein superfamilies: analysis of SH2 domains. Proc Int Conf Intell Syst Mol Biol. 1998; 6:165-74.