TreeDet
TreeDet predicts subfamily-specific functional residues in protein sequence alignments by integrating tree-based phylogenetic, correlation-based distance-matrix, and PCA/clustering analyses to identify conserved positions relevant to protein function.
Key Features:
- Integration of multiple analytical methods: Combines tree-based phylogenetic representations, correlation-based distance-matrix comparisons, and PCA with clustering to detect conserved and subfamily-specific positions.
- Tree-based phylogenetic analysis: Uses phylogenetic representations to relate sequence conservation to evolutionary context.
- Correlation-based distance-matrix comparisons: Compares mutational behaviors using distance matrices to detect correlated changes among positions.
- PCA with clustering: Applies principal component analysis coupled with clustering to analyze sequence distributions in multidimensional sequence space.
- Automatic subfamily partitioning using Relative Entropy: Determines optimal subfamily divisions based on Information Theory measures, specifically Relative Entropy.
- Comparison of mutational behaviors: Compares mutational behaviors of full-length proteins with those of specific sequence positions to identify functional specificity.
- Sequence-space vector representation and dimensionality reduction: Represents proteins and residues as vectors in a generalized sequence space and projects them into lower-dimensional spaces to identify residue groups specific to subfamilies.
- Statistical validation on curated families: Validated across non-redundant protein family lists, including ligand-binding families and those with annotated functional sites, demonstrating prediction of residues near bound ligands or key functional amino acids.
- Reliability assessment of positions: Provides a reliability assessment that highlights the confidence of alignment positions used for predictions.
Scientific Applications:
- Identification of functional residues: Identifying residues critical to protein activity and specificity within protein families.
- Structural biology: Mapping predicted functional residues onto structures to interpret ligand interactions and functional sites.
- Molecular evolution studies: Investigating evolutionary patterns of conservation and divergence across subfamilies.
- Drug design: Prioritizing subfamily-specific residues that may influence ligand binding or selectivity for therapeutic targeting.
- Experimental hypothesis generation: Generating testable hypotheses about residues to validate by biochemical or mutational experiments.
Methodology:
Analyzes multiple sequence alignments using tree-based phylogenetic representations, correlation-based distance-matrix comparisons of mutational behaviors, PCA with clustering of sequence distributions, automatic subfamily partitioning via Relative Entropy, vector representation of proteins and residues in sequence space with projection to lower-dimensional spaces, comparison of full-length versus position-specific mutational behaviors, statistical validation against non-redundant protein family lists including ligand-binding and annotated functional sites, and assessment of positional reliability.
Topics
Details
- Tool Type:
- web application
- Programming Languages:
- PHP, Perl, Fortran, C
- Added:
- 2/10/2017
- Last Updated:
- 11/25/2024
Operations
Publications
Carro A, Tress M, de Juan D, Pazos F, Lopez-Romero P, del Sol A, Valencia A, Rojas AM. TreeDet: a web server to explore sequence space. Nucleic Acids Research. 2006;34(Web Server):W110-W115. doi:10.1093/nar/gkl203. PMID:16844971. PMCID:PMC1538789.
del Sol Mesa A, Pazos F, Valencia A. Automatic Methods for Predicting Functionally Important Residues. Journal of Molecular Biology. 2003;326(4):1289-1302. doi:10.1016/s0022-2836(02)01451-1. PMID:12589769.
Casari G, Sander C, Valencia A. A method to predict functional residues in proteins. Nature Structural & Molecular Biology. 1995;2(2):171-178. doi:10.1038/nsb0295-171. PMID:7749921.
Tress ML, Graña O, Valencia A. SQUARE—determining reliable regions in sequence alignments. Bioinformatics. 2004;20(6):974-975. doi:10.1093/bioinformatics/bth032. PMID:14764569.