Text to Image Generation with Semantic-Spatial Aware GAN
Wentong Liao, Kai Hu, Michael Ying Yang, Bodo Rosenhahn
摘要
A text to image generation (T2I) model aims to generate photo-realistic images which are semantically consistent with the text descriptions. Built upon the recent advances in generative adversarial networks (GANs), existing T2I models have made great progress. However, a close inspection of their generated images reveals two major limitations: (1) The condition batch normalization methods are applied on the whole image feature maps equally, ignoring the local semantics; (2) The text encoder is fixed during training, which should be trained with the image generator jointly to learn better text representations for image generation. To address these limitations, we propose a novel framework Semantic-Spatial Aware GAN, which is trained in an end-to-end fashion so that the text encoder can exploit better text information. Concretely, we introduce a novel Semantic-Spatial Aware Convolution Network, which (1) learns semantic-adaptive transformation conditioned on text to effectively fuse text features and image features, and (2) learns a mask map in a weakly-supervised way that depends on the current text-image fusion process in order to guide the transformation spatially. Experiments on the challenging COCO and CUB bird datasets demonstrate the advantage of our method over the recent state-of-the-art approaches, regarding both visual fidelity and alignment with input text description. * Kai Hu contributed to this work when he was doing his master thesis supervised by Wentong Liao in TNT. † Equal contribution.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper20
- TextDiffuser: Diffusion Models as Text PaintersJingye Chen, Yupan Huang, Tengchao Lv, Lei Cui 等NeurIPS 2023 · 被引用 290 次
- Text-IF: Leveraging Semantic Text Guidance for Degradation-Aware and Interactive Image FusionXunpeng Yi, Han Xu, Hao Zhang, Linfeng Tang 等CVPR 2024 · 被引用 121 次
- Frido: Feature Pyramid Diffusion for Complex Scene Image SynthesisWan-Cyuan Fan, Yen-Chun Chen, Dongdong Chen, Yu Cheng 等AAAI 2023 · 被引用 118 次
- Social Reward: Evaluating and Enhancing Generative AI through Million-User Feedback from an Online Creative CommunityArman Isajanyan, Artur Shatveryan, David Kocharian, Zhangyang Wang 等ICLR 2024 · 被引用 9 次
- Capability-aware Prompt Reformulation Learning for Text-to-Image GenerationJingtao Zhan, Qingyao Ai, Yiqun Liu, Jia Chen 等SIGIR 2024 · 被引用 7 次
它引用的顶会 Paper2
相关 Paper
- Semantics-Enhanced Adversarial Nets for Text-to-Image SynthesisHongchen Tan, Xiuping Liu, Xin Li, Yi Zhang 等ICCV 2019 · 被引用 80 次
- Background Layout Generation and Object Knowledge Transfer for Text-to-Image GenerationZhuowei Chen, Zhendong Mao, Shancheng Fang, Bo HuACM MM 2022 · 被引用 6 次
- Adma-GAN: Attribute-Driven Memory Augmented GANs for Text-to-Image GenerationXintian Wu, Hanbin Zhao, Liangli Zheng, Shouhong Ding 等ACM MM 2022 · 被引用 17 次
- DSE-GAN: Dynamic Semantic Evolution Generative Adversarial Network for Text-to-Image GenerationMengqi Huang, Zhendong Mao, Penghui Wang, Quan Wang 等ACM MM 2022 · 被引用 26 次
- DAE-GAN: Dynamic Aspect-aware GAN for Text-to-Image SynthesisShulan Ruan, Yong Zhang, Kun Zhang, Yanbo Fan 等ICCV 2021 · 被引用 125 次
