ClusterBurden

ClusterBurden integrates gene burden and amino acid residue position to improve prediction of gene- and missense rare-variant pathogenicity in case-control studies of protein-coding regions such as Mendelian diseases.


Key Features:

  • Joint Analysis Framework: integrates gene burden with amino acid residue position to jointly evaluate gene- and missense-variant pathogenicity.
  • BIN-test for Clustering Detection: implements the BIN-test to detect missense variant clustering within protein regions, with simulations showing higher power than Anderson-Darling and Kolmogorov-Smirnov tests.
  • Generalized Additive Modeling (GAM): uses GAMs to identify mutational hotspots by integrating burden and clustering information and to refine regional burden maps.
  • Hotspot Models: provides hotspot and hotspot+ models that leverage burden, clustering information, and in silico predictors to assess mutational hotspots.
  • Performance and Efficiency: demonstrated increased computational efficiency and statistical power in analyses of hypertrophic cardiomyopathy cohorts, in some cases equivalent to increasing sample size by 50% for genes with strong clustering signals.
  • Integration with ACMG Criteria: hotspot+ models integrate ACMG criteria PM1 and PP3 to provide evidence for pathogenicity and support reclassification of variants of uncertain significance.

Scientific Applications:

  • Mendelian disease gene discovery: increases power to detect disease-associated genes in case-control studies of rare missense variants that cluster within protein regions.
  • Variant interpretation and reclassification: refines pathogenicity assessments by combining positional, burden, and in silico evidence and supports application of ACMG PM1 and PP3.
  • Hypertrophic cardiomyopathy analyses: applied to HCM cohorts to identify gene-specific clustering and mutational hotspots.

Methodology:

Performs a joint analysis of gene burden and amino acid residue position; applies the BIN-test to detect missense clustering and compares its power to Anderson-Darling and Kolmogorov-Smirnov in simulations; fits generalized additive models to integrate burden and clustering and to define hotspot and hotspot+ models that incorporate in silico predictors and integrate ACMG PM1 and PP3.

Topics

Details

Tool Type:
library
Programming Languages:
R
Added:
1/14/2020
Last Updated:
11/24/2024

Operations

Publications

Waring A, Harper A, Salatino S, Kramer C, Neubauer S, Thomson K, Watkins H, Farrall M. Data-driven modelling of mutational hotspots and <i>in-silico</i> predictors in hypertrophic cardiomyopathy. Unknown Journal. 2019. doi:10.1101/826164.

Waring A, Harper A, Salatino S, Kramer C, Neubauer S, Thomson K, Watkins H, Farrall M. Data-driven modelling of mutational hotspots and in silico predictors in hypertrophic cardiomyopathy. Journal of Medical Genetics. 2020;58(8):556-564. doi:10.1136/jmedgenet-2020-106922. PMID:32732227. PMCID:PMC8327322.

PMID: 32732227
PMCID: PMC8327322
Funding: - British Heart Foundation: RE/13/1/30181 - National Heart, Lung, and Blood Institute: U01HL117006-01A1 - Wellcome Trust: 203141/Z/16/Z, 203834/Z/16/Z