sugsvarsel

sugsvarsel implements variable selection and fast approximate Bayesian inference in Dirichlet process (DP) mixture models for model-based clustering of high-dimensional biological data.


Key Features:

  • SUGS algorithm: Implements the Sequential Updating and Greedy Search (SUGS) algorithm for fast approximate Bayesian inference in DP mixture models.
  • SUGS VarSel extension: Extends SUGS to perform variable selection within the DP mixture model framework.
  • Variable selection: Identifies relevant variables/features to improve interpretability and clustering accuracy in high-dimensional datasets.
  • Bayesian model selection (BMS): Poses clustering as a Bayesian model selection task to compare alternative models.
  • Bayesian model averaging (BMA): Supports Bayesian model averaging to combine inference across multiple models.
  • Approximate inference without MCMC: Infers cluster structure and number of clusters without relying on Markov chain Monte Carlo (MCMC) methods.
  • Data applicability: Targets high-dimensional biological datasets including pan-cancer proteomics, reverse-phase protein array (RPPA) data, TCGA datasets, and cancer transcriptomics.
  • Implementation and performance: Implemented in R with C++ computational kernels and automated parallel processing for computational efficiency, demonstrated on large-scale TCGA proteomic datasets.

Scientific Applications:

  • Pan-cancer proteomic characterization: Clustering and variable selection applied to pan-cancer proteomics analyses of thousands of TCGA tumor samples.
  • RPPA data analysis: Analysis of reverse-phase protein array (RPPA) datasets from TCGA for identification of relevant protein biomarkers.
  • Cancer transcriptomics: Evaluation and application to cancer transcriptomics datasets in simulation studies and empirical analyses.
  • Model-based clustering and biomarker discovery: General use for model-based clustering and selection of informative features in high-dimensional biomedical data.

Methodology:

Implements the Sequential Updating and Greedy Search (SUGS) algorithm and the SUGS VarSel extension for approximate Bayesian inference in Dirichlet process mixture models, frames clustering as Bayesian model selection with support for Bayesian model averaging, and uses an R implementation with C++ kernels and automated parallel processing.

Topics

Details

Programming Languages:
R, C++
Added:
1/14/2020
Last Updated:
11/24/2024

Operations

Publications

Crook OM, Gatto L, Kirk PDW. Fast approximate inference for variable selection in Dirichlet process mixtures, with an application to pan-cancer proteomics. Statistical Applications in Genetics and Molecular Biology. 2019;18(6). doi:10.1515/sagmb-2018-0065. PMID:31829970. PMCID:PMC7614016.

PMID: 31829970
PMCID: PMC7614016
Funding: - Medical Research Council: MC_UU_00002/10