DR-GAN

DR-GAN generates images from textual descriptions by applying distribution regularization to improve fidelity and semantic accuracy in text-to-image (T2I) generation.


Key Features:

  • Text-to-Image (T2I) generation: Performs image synthesis from textual descriptions using an adversarial learning framework.
  • Semantic Disentangling Module (SDM): Uses a Spatial Self-Attention Mechanism (SSAM) and a Semantic Disentangling Loss (SDL) to distill and emphasize key semantic elements from text for the generator.
  • Distribution Normalization Module (DNM): Employs a variational auto-encoder (VAE) to normalize and denoise latent image distributions and refine latent representations.
  • Distribution Adversarial Loss (DAL): Guides the generator to align its latent outputs with normalized real image distributions in latent space.
  • Generative Adversarial Framework: Trains a generator and discriminator adversarially, with the discriminator evaluating synthesized images against real images.

Scientific Applications:

  • Data Augmentation: Generating diverse synthetic visual examples from textual descriptions to augment training datasets where real data are limited.
  • Creative Design and Art Generation: Translating textual concepts into visual forms to support generative design and artistic exploration.
  • Medical Imaging: Producing synthetic medical images from textual descriptions of conditions to support research and development when annotated data are scarce.

Methodology:

The model uses an adversarial framework in which the generator creates images from text inputs guided by the SDM (SSAM and SDL) and the DNM (VAE and DAL), while the discriminator evaluates synthesized images against real images and adversarial losses update the networks.

Topics

Details

License:
MIT
Cost:
Free of charge
Tool Type:
command-line tool
Operating Systems:
Linux
Programming Languages:
Python
Added:
7/26/2022
Last Updated:
11/24/2024

Operations

Publications

Tan H, Liu X, Yin B, Li X. DR-GAN: Distribution Regularization for Text-to-Image Generation. IEEE Transactions on Neural Networks and Learning Systems. 2023;34(12):10309-10323. doi:10.1109/tnnls.2022.3165573. PMID:35442894.

PMID: 35442894
Funding: - National Key Research and Development Program of China: 2021ZD0111900 - National Natural Science Foundation of China: 61976040 - National Science Foundation of USA: CBET-2115405, OIA-1946231 - Chinese Postdoctoral Science Foundation: 2021M700303