HAC

HAC performs hierarchical agglomerative clustering using a hierarchical stochastic block model to infer and summarize multilevel structure in heterogeneous biological interaction networks for clustering, link prediction, and functional annotation.


Key Features:

  • Hierarchical Stochastic Block Model: Employs a hierarchical stochastic block model to infer multilevel network structure.
  • Maximum Likelihood Estimation: Model inference is driven by maximum likelihood estimation to estimate block assignments.
  • Bayesian Model Selection: Utilizes Bayesian model selection to collapse fine-structured groups and identify top-level clusters.
  • Simultaneous Analysis of Multiple Interaction Types: Scores models additively over independent interaction types to enable joint analysis of diverse interactions.
  • Link Prediction and Cross-Validation: Generates link predictions and validates them using cross-validation to quantify prediction performance.
  • Comparative Performance: Demonstrates superior link-prediction performance relative to other clustering algorithms and model-free graph diffusion kernels on genome-scale datasets.

Scientific Applications:

  • Large-scale biological network analysis: Applied to genome-scale interaction networks to summarize structure and predict unobserved interactions.
  • Protein-protein interaction analysis: Used on genome-scale yeast protein interactions to identify ~100 top-level clusters and ~1,000 fine-level clusters with an average of ~5 proteins each.
  • Functional annotation and inference: Facilitates assignment of functional annotations and assessment of molecular complexes from clustered interaction data.
  • Assessment of joint physical and genetic data: Reveals that joint clustering of physical and genetic interaction data may not always yield synergistic results in current high-throughput datasets.

Methodology:

Inference uses a hierarchical stochastic block model fitted by maximum likelihood estimation, Bayesian model selection to collapse groups and identify top-level clusters, additive scoring across independent interaction types, and cross-validation to validate link predictions.

Topics

Details

Tool Type:
command-line tool
Operating Systems:
Linux
Added:
8/3/2017
Last Updated:
11/25/2024

Operations

Publications

Park Y, Bader JS. Resolving the structure of interactomes with hierarchical agglomerative clustering. BMC Bioinformatics. 2011;12(S1). doi:10.1186/1471-2105-12-s1-s44. PMID:21342576. PMCID:PMC3044301.

Links