HiFiX

HiFiX performs high-fidelity clustering of homologous protein sequences to separate highly divergent homologs and resolve families affected by protein domain rearrangements.


Key Features:

  • Comprehensive Clustering Strategy: Clusters sequences based on entire-sequence length and gene-family-specific substitution patterns to increase sensitivity and accuracy.
  • Pairwise Similarity and Pre-family Formation: Performs pairwise sequence comparisons using permissive similarity criteria to group sequences into pre-families.
  • Homogeneous Cluster Refinement: Subdivides pre-families by analyzing the topology of the similarity network to obtain more homogeneous clusters.
  • Family Merging and Model Selection: Computes multiple sequence alignments and applies model selection techniques to merge clusters into statistically supported families.
  • Robustness to Sequence Divergence: Adapts to sequence divergence and protein domain rearrangements to distinguish highly divergent homologs from partial homologs.
  • Scalability: Designed to operate on very large datasets while maintaining clustering accuracy and computational feasibility.

Scientific Applications:

  • Genomics: Delineates gene families to support analyses of gene content and family-level evolution across genomes.
  • Proteomics: Facilitates grouping of protein sequences for studies of protein function and functional divergence.
  • Phylogenetics: Supports inference of phylogenetic relationships and the functional evolution of genes through robust family definitions.
  • Protein domain architecture analysis: Identifies and separates sequences affected by domain rearrangements and partial homology for domain-centric studies.

Methodology:

Performs pairwise sequence comparisons with permissive similarity criteria to form pre-families; analyzes similarity network topology to subdivide pre-families into homogeneous clusters; computes multiple sequence alignments and applies model selection to merge clusters into final families.

Topics

Details

Tool Type:
command-line tool
Operating Systems:
Linux, Mac
Programming Languages:
Python
Added:
12/18/2017
Last Updated:
12/10/2018

Operations

Publications

Miele V, Penel S, Daubin V, Picard F, Kahn D, Duret L. High-quality sequence clustering guided by network topology and multiple alignment likelihood. Bioinformatics. 2012;28(8):1078-1085. doi:10.1093/bioinformatics/bts098.

Documentation

Links