HiFiX
HiFiX performs high-fidelity clustering of homologous protein sequences to separate highly divergent homologs and resolve families affected by protein domain rearrangements.
Key Features:
- Comprehensive Clustering Strategy: Clusters sequences based on entire-sequence length and gene-family-specific substitution patterns to increase sensitivity and accuracy.
- Pairwise Similarity and Pre-family Formation: Performs pairwise sequence comparisons using permissive similarity criteria to group sequences into pre-families.
- Homogeneous Cluster Refinement: Subdivides pre-families by analyzing the topology of the similarity network to obtain more homogeneous clusters.
- Family Merging and Model Selection: Computes multiple sequence alignments and applies model selection techniques to merge clusters into statistically supported families.
- Robustness to Sequence Divergence: Adapts to sequence divergence and protein domain rearrangements to distinguish highly divergent homologs from partial homologs.
- Scalability: Designed to operate on very large datasets while maintaining clustering accuracy and computational feasibility.
Scientific Applications:
- Genomics: Delineates gene families to support analyses of gene content and family-level evolution across genomes.
- Proteomics: Facilitates grouping of protein sequences for studies of protein function and functional divergence.
- Phylogenetics: Supports inference of phylogenetic relationships and the functional evolution of genes through robust family definitions.
- Protein domain architecture analysis: Identifies and separates sequences affected by domain rearrangements and partial homology for domain-centric studies.
Methodology:
Performs pairwise sequence comparisons with permissive similarity criteria to form pre-families; analyzes similarity network topology to subdivide pre-families into homogeneous clusters; computes multiple sequence alignments and applies model selection to merge clusters into final families.
Topics
Details
- Tool Type:
- command-line tool
- Operating Systems:
- Linux, Mac
- Programming Languages:
- Python
- Added:
- 12/18/2017
- Last Updated:
- 12/10/2018
Operations
Publications
Miele V, Penel S, Daubin V, Picard F, Kahn D, Duret L. High-quality sequence clustering guided by network topology and multiple alignment likelihood. Bioinformatics. 2012;28(8):1078-1085. doi:10.1093/bioinformatics/bts098.