AD-AE

AD-AE disentangles confounding factors from gene expression data to produce biologically informative, transferable latent embeddings that preserve true biological signals while excluding technical and non-informative biological variation.


Key Features:

  • Autoencoder: Generates embeddings that reconstruct the original gene expression measurements.
  • Adversary Network: Predicts confounders from the embeddings to drive adversarial deconfounding.
  • Adversarial training: Jointly optimizes the autoencoder and adversary to minimize confounder encoding in the latent space.
  • Deconfounding scope: Targets technical artifacts (e.g., batch effects) and non-informative biological variables (e.g., age).
  • Preservation of biology: Retains genuine biological signals within the embeddings despite deconfounding.
  • Transferability: Produces embeddings that generalize across domains with varying confounder distributions.
  • Deep unsupervised architecture: Employs deep unsupervised neural networks to extract meaningful latent spaces from gene expression profiles.
  • Empirical performance: Demonstrated on two distinct gene expression datasets and shown to outperform standard autoencoders and other deconfounding methods.

Scientific Applications:

  • Cross-dataset analysis: Enables analyses across datasets by removing confounder-driven variation that hinders generalization.
  • Dataset integration: Facilitates integration of diverse gene expression datasets without interference from technical or non-informative biological confounders.
  • Transferable latent representations: Produces embeddings suitable for comparative studies that require representations that generalize across differing confounder distributions.

Methodology:

AD-AE trains an autoencoder to reconstruct gene expression and an adversary network to predict confounders from the embeddings, using adversarial training to jointly optimize the networks to remove confounder information while preserving biological signal.

Topics

Details

Tool Type:
command-line tool
Programming Languages:
Python
Added:
3/19/2021
Last Updated:
4/11/2021

Operations

Publications

Dincer AB, Janizek JD, Lee S. Adversarial deconfounding autoencoder for learning robust gene expression embeddings. Bioinformatics. 2020;36(Supplement_2):i573-i582. doi:10.1093/bioinformatics/btaa796. PMID:33381842. PMCID:PMC7773484.

PMID: 33381842
PMCID: PMC7773484
Funding: - National Institutes of Health: R01 NIA AG 061132, R35 GM 128638 - CAREER: DBI-1552309, DBI-1759487