HyDRA
HyDRA integrates score-based and distance-based rank aggregation to prioritize candidate disease-associated genes by combining multiple similarity criteria and weighting top-ranked candidates.
Key Features:
- Hybrid aggregation methodology: HyDRA combines score-based and distance-based rank aggregation approaches to leverage complementary strengths of each method.
- Top-versus-bottom (TvB) weighting: TvB weighting prioritizes reliability at the top of gene rankings to reflect the focus on highest-ranked candidates for experimental validation.
- Predictive quality assessment and cross-validation: The system evaluates predictive quality of aggregation methods, including approaches derived from machine learning and social choice theory, and conducts cross-validation to assess the impact of the number of training genes and similarity criteria.
- Iterative gene discovery: An iterative procedure augments the set of training genes with newly discovered genes across rounds to improve consistency and robustness of prioritization.
- Performance and best practices: Methods implemented in HyDRA were reported to outperform ToppGene and Endeavour and recommend using the union of top-ranked items from different aggregation methods for final prioritization.
Scientific Applications:
- Iterative gene discovery in tumor types: Applied to iterative gene discovery in glioblastoma, meningioma, and breast cancer using sequentially augmented training gene lists.
- Syndrome-associated gene identification: Used sequentially augmented lists related to Turcot syndrome and Li-Fraumeni condition to facilitate identification of novel disease-associated genes.
- Benchmarking across disease cohorts: Tested on gene sets for autism, breast cancer, colorectal cancer, endometriosis, ischaemic stroke, leukemia, lymphoma, and osteoarthritis to evaluate aggregation performance.
Methodology:
Hybrid aggregation combining score-based and distance-based approaches; top-versus-bottom (TvB) weighting; methods and properties drawn from social choice theory, political science, computer science, and statistics; techniques derived from machine learning; in-depth cross-validation studies evaluating the number of training genes and similarity criteria; and an iterative augmentation procedure for training gene sets.
Topics
Details
- Tool Type:
- command-line tool
- Operating Systems:
- Linux, Windows, Mac
- Programming Languages:
- MATLAB
- Added:
- 8/3/2017
- Last Updated:
- 11/25/2024
Operations
Publications
Kim M, Farnoud F, Milenkovic O. HyDRA: gene prioritization via hybrid distance-score rank aggregation. Bioinformatics. 2014;31(7):1034-1043. doi:10.1093/bioinformatics/btu766. PMID:25411330.