Text to Image Generation with Semantic-Spatial Aware GAN
Wentong Liao, Kai Hu, Michael Ying Yang, Bodo Rosenhahn
Abstract
A text to image generation (T2I) model aims to generate photo-realistic images which are semantically consistent with the text descriptions. Built upon the recent advances in generative adversarial networks (GANs), existing T2I models have made great progress. However, a close inspection of their generated images reveals two major limitations: (1) The condition batch normalization methods are applied on the whole image feature maps equally, ignoring the local semantics; (2) The text encoder is fixed during training, which should be trained with the image generator jointly to learn better text representations for image generation. To address these limitations, we propose a novel framework Semantic-Spatial Aware GAN, which is trained in an end-to-end fashion so that the text encoder can exploit better text information. Concretely, we introduce a novel Semantic-Spatial Aware Convolution Network, which (1) learns semantic-adaptive transformation conditioned on text to effectively fuse text features and image features, and (2) learns a mask map in a weakly-supervised way that depends on the current text-image fusion process in order to guide the transformation spatially. Experiments on the challenging COCO and CUB bird datasets demonstrate the advantage of our method over the recent state-of-the-art approaches, regarding both visual fidelity and alignment with input text description. * Kai Hu contributed to this work when he was doing his master thesis supervised by Wentong Liao in TNT. † Equal contribution.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers20
- TextDiffuser: Diffusion Models as Text PaintersJingye Chen, Yupan Huang, Tengchao Lv, Lei Cui et al.NeurIPS 2023 · 290 citations
- Text-IF: Leveraging Semantic Text Guidance for Degradation-Aware and Interactive Image FusionXunpeng Yi, Han Xu, Hao Zhang, Linfeng Tang et al.CVPR 2024 · 121 citations
- Frido: Feature Pyramid Diffusion for Complex Scene Image SynthesisWan-Cyuan Fan, Yen-Chun Chen, Dongdong Chen, Yu Cheng et al.AAAI 2023 · 118 citations
- Social Reward: Evaluating and Enhancing Generative AI through Million-User Feedback from an Online Creative CommunityArman Isajanyan, Artur Shatveryan, David Kocharian, Zhangyang Wang et al.ICLR 2024 · 9 citations
- Capability-aware Prompt Reformulation Learning for Text-to-Image GenerationJingtao Zhan, Qingyao Ai, Yiqun Liu, Jia Chen et al.SIGIR 2024 · 7 citations
Builds on2
Related papers
- Semantics-Enhanced Adversarial Nets for Text-to-Image SynthesisHongchen Tan, Xiuping Liu, Xin Li, Yi Zhang et al.ICCV 2019 · 80 citations
- Background Layout Generation and Object Knowledge Transfer for Text-to-Image GenerationZhuowei Chen, Zhendong Mao, Shancheng Fang, Bo HuACM MM 2022 · 6 citations
- Adma-GAN: Attribute-Driven Memory Augmented GANs for Text-to-Image GenerationXintian Wu, Hanbin Zhao, Liangli Zheng, Shouhong Ding et al.ACM MM 2022 · 17 citations
- DSE-GAN: Dynamic Semantic Evolution Generative Adversarial Network for Text-to-Image GenerationMengqi Huang, Zhendong Mao, Penghui Wang, Quan Wang et al.ACM MM 2022 · 26 citations
- DAE-GAN: Dynamic Aspect-aware GAN for Text-to-Image SynthesisShulan Ruan, Yong Zhang, Kun Zhang, Yanbo Fan et al.ICCV 2021 · 125 citations
