Variational Hetero-Encoder Randomized GANs for Joint Image-Text Modeling
Hao Zhang, Bo Chen, Long Tian, Zhengjue Wang, Mingyuan Zhou
Abstract
For bidirectional joint image-text modeling, we develop variational hetero-encoder (VHE) randomized generative adversarial network (GAN), a versatile deep generative model that integrates a probabilistic text decoder, probabilistic image encoder, and GAN into a coherent end-to-end multi-modality learning framework. VHE randomized GAN (VHE-GAN) encodes an image to decode its associated text, and feeds the variational posterior as the source of randomness into the GAN image generator. We plug three off-the-shelf modules, including a deep topic model, a ladder-structured image encoder, and StackGAN++, into VHE-GAN, which already achieves competitive performance. This further motivates the development of VHE-raster-scan-GAN that generates photo-realistic images in not only a multi-scale low-to-high-resolution manner, but also a hierarchical-semantic coarse-to-fine fashion. By capturing and relating hierarchical semantic and visual concepts with end-to-end training, VHE-raster-scan-GAN achieves state-of-the-art performance in a wide variety of image-text multi-modality learning and generation tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- Diffusion-GAN: Training GANs with DiffusionZhendong Wang, Huangjie Zheng, Pengcheng He, Weizhu Chen et al.ICLR 2023 · 65 citations
- Sawtooth Factorial Topic Embeddings Guided Gamma Belief NetworkZhibin Duan, Dongsheng Wang, Bo Chen, Chaojie Wang et al.ICML 2021 · 49 citations
- Exploiting Chain Rule and Bayes' Theorem to Compare Probability DistributionsHuangjie Zheng, Mingyuan ZhouNeurIPS 2021 · 33 citations
- Thompson Sampling via Local UncertaintyZhendong Wang, Mingyuan ZhouICML 2020 · 21 citations
- Alleviating "Posterior Collapse" in Deep Topic Models via Policy GradientYewen Li, Chaojie Wang, Zhibin Duan, Dongsheng Wang et al.NeurIPS 2022 · 11 citations
Related papers
- Recurrent Hierarchical Topic-Guided RNN for Language GenerationDandan Guo, Bo Chen, Ruiying Lu, Mingyuan ZhouICML 2020 · 20 citations
- L-Verse: Bidirectional Generation Between Image and TextTaehoon Kim, Gwangmo Song, Sihaeng Lee, Sangyun Kim et al.CVPR 2022 · 23 citations
- DF-GAN: A Simple and Effective Baseline for Text-to-Image SynthesisMing Tao, Hao Tang, Fei Wu, Xiaoyuan Jing et al.CVPR 2022 · 296 citations
- Semantics-Enhanced Adversarial Nets for Text-to-Image SynthesisHongchen Tan, Xiuping Liu, Xin Li, Yi Zhang et al.ICCV 2019 · 80 citations
- TediGAN: Text-Guided Diverse Face Image Generation and ManipulationWeihao Xia, Yujiu Yang, Jing-Hao Xue, Baoyuan WuCVPR 2021
