mixedCCA

mixedCCA performs sparse canonical correlation analysis (CCA) using latent Gaussian copula models to estimate dependencies between continuous, binary, and zero-inflated variables for multi-view data integration and is implemented as an R package.


Key Features:

  • Integration of mixed variable types: Uses latent Gaussian copula models to represent dependencies among continuous, binary, and zero-inflated variables.
  • Efficient computational approach: Estimates latent correlations with a hybrid multilinear interpolation and optimization scheme to accelerate computation.
  • Scalability and performance: Applies to high-dimensional settings where the number of variables p is large, reports speed-ups of several orders of magnitude relative to existing methods, and provides theoretical guarantees on approximation error.
  • Demonstrated applications: Validated on simulated data and applied to high-dimensional sparse quantitative and relative abundance microbiome data and multi-view datasets from The Cancer Genome Atlas Project.

Scientific Applications:

  • Microbiome research: Analysis and integration of high-dimensional, sparse quantitative and relative abundance microbiome data comprising mixed variable types.
  • Cancer genomics: Integration and exploration of multi-view genomic datasets from The Cancer Genome Atlas Project to study relationships across data types.

Methodology:

Estimation of latent Gaussian correlations via a hybrid multilinear interpolation and optimization approach to perform sparse CCA, with theoretical approximation-error guarantees.

Topics

Details

License:
Not licensed
Tool Type:
library
Operating Systems:
Mac, Linux, Windows
Programming Languages:
R, C++
Added:
6/28/2022
Last Updated:
11/24/2024

Operations

Publications

Yoon G, Müller CL, Gaynanova I. Fast Computation of Latent Correlations. Journal of Computational and Graphical Statistics. 2021;30(4):1249-1256. doi:10.1080/10618600.2021.1882468. PMID:35280976. PMCID:PMC8916743.

PMID: 35280976
PMCID: PMC8916743
Funding: - National Institutes of Health National Cancer Institute: T32-CA090301 - National Science Foundation: DMS-1712943