HyDRA

HyDRA integrates score-based and distance-based rank aggregation to prioritize candidate disease-associated genes by combining multiple similarity criteria and weighting top-ranked candidates.


Key Features:

  • Hybrid aggregation methodology: HyDRA combines score-based and distance-based rank aggregation approaches to leverage complementary strengths of each method.
  • Top-versus-bottom (TvB) weighting: TvB weighting prioritizes reliability at the top of gene rankings to reflect the focus on highest-ranked candidates for experimental validation.
  • Predictive quality assessment and cross-validation: The system evaluates predictive quality of aggregation methods, including approaches derived from machine learning and social choice theory, and conducts cross-validation to assess the impact of the number of training genes and similarity criteria.
  • Iterative gene discovery: An iterative procedure augments the set of training genes with newly discovered genes across rounds to improve consistency and robustness of prioritization.
  • Performance and best practices: Methods implemented in HyDRA were reported to outperform ToppGene and Endeavour and recommend using the union of top-ranked items from different aggregation methods for final prioritization.

Scientific Applications:

  • Iterative gene discovery in tumor types: Applied to iterative gene discovery in glioblastoma, meningioma, and breast cancer using sequentially augmented training gene lists.
  • Syndrome-associated gene identification: Used sequentially augmented lists related to Turcot syndrome and Li-Fraumeni condition to facilitate identification of novel disease-associated genes.
  • Benchmarking across disease cohorts: Tested on gene sets for autism, breast cancer, colorectal cancer, endometriosis, ischaemic stroke, leukemia, lymphoma, and osteoarthritis to evaluate aggregation performance.

Methodology:

Hybrid aggregation combining score-based and distance-based approaches; top-versus-bottom (TvB) weighting; methods and properties drawn from social choice theory, political science, computer science, and statistics; techniques derived from machine learning; in-depth cross-validation studies evaluating the number of training genes and similarity criteria; and an iterative augmentation procedure for training gene sets.

Topics

Details

Tool Type:
command-line tool
Operating Systems:
Linux, Windows, Mac
Programming Languages:
MATLAB
Added:
8/3/2017
Last Updated:
11/25/2024

Operations

Publications

Kim M, Farnoud F, Milenkovic O. HyDRA: gene prioritization via hybrid distance-score rank aggregation. Bioinformatics. 2014;31(7):1034-1043. doi:10.1093/bioinformatics/btu766. PMID:25411330.

Documentation

Links