DAE-GAN: Dynamic Aspect-aware GAN for Text-to-Image Synthesis
Shulan Ruan, Yong Zhang, Kun Zhang, Yanbo Fan, Fan Tang, Qi Liu, Enhong Chen
Abstract
Text-to-image synthesis refers to generating an image from a given text description, the key goal of which lies in photo realism and semantic consistency. Previous methods usually generate an initial image with sentence embedding and then refine it with fine-grained word embedding. Despite the significant progress, the ‘aspect’ information (e.g., red eyes) contained in the text, referring to several words rather than a word that depicts ‘a particular part or feature of something’, is often ignored, which is highly helpful for synthesizing image details. How to make better utilization of aspect information in text-to-image synthesis still remains an unresolved challenge. To address this problem, in this paper, we propose a Dynamic Aspect-awarE GAN (DAE-GAN) that represents text information comprehensively from multiple granularities, including sentence-level, word-level, and aspect-level. Moreover, inspired by human learning behaviors, we develop a novel Aspect-aware Dynamic Re-drawer (ADR) for image refinement, in which an Attended Global Refinement (AGR) module and an Aspect-aware Local Refinement (ALR) module are alternately employed. AGR utilizes word-level embedding to globally enhance the previously generated image, while ALR dynamically employs aspect-level embedding to refine image details from a local perspective. Finally, a corresponding matching loss function is designed to ensure the text-image semantic consistency at different levels. Extensive experiments on two well-studied and publicly available datasets (i.e., CUB-200 and COCO) demonstrate the superiority and rationality of our method.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers18
- Vector Quantized Diffusion Model for Text-to-Image SynthesisShuyang Gu, Dong Chen, Jianmin Bao, Fang Wen et al.CVPR 2022 · 607 citations
- DF-GAN: A Simple and Effective Baseline for Text-to-Image SynthesisMing Tao, Hao Tang, Fei Wu, Xiaoyuan Jing et al.CVPR 2022 · 296 citations
- Draw Your Art Dream: Diverse Digital Art Synthesis with Multimodal Guided DiffusionNisha Huang, Fan Tang, Weiming Dong, Changsheng XuACM MM 2022 · 49 citations
- X-Mesh: Towards Fast and Accurate Text-driven 3D Stylization via Dynamic Textual GuidanceYiwei Ma, Haowei Wang, Xiaoqing Zhang, Guannan Jiang et al.ICCV 2023 · 48 citations
- StyleT2I: Toward Compositional and High-Fidelity Text-to-Image SynthesisZhiheng Li, Martin Renqiang Min, Kai Li, Chenliang XuCVPR 2022 · 38 citations
Builds on7
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
- Knowing What, How and Why: A Near Complete Solution for Aspect-Based Sentiment AnalysisHaiyun Peng, Lu Xu, Lidong Bing, Fei Huang et al.AAAI 2020 · 494 citations
- Relation-Aware Collaborative Learning for Unified Aspect-Based Sentiment AnalysisZhuang Chen, Tieyun QianACL 2020 · 194 citations
- Image Synthesis From Reconfigurable Layout and StyleWei Sun, Tianfu WuICCV 2019 · 160 citations
- Replicate, Walk, and Stop on Syntax: An Effective Neural Network Model for Aspect-Level Sentiment ClassificationYaowei Zheng, Richong Zhang, Samuel Mensah, Yongyi MaoAAAI 2020 · 48 citations
Related papers
- Semantics-Enhanced Adversarial Nets for Text-to-Image SynthesisHongchen Tan, Xiuping Liu, Xin Li, Yi Zhang et al.ICCV 2019 · 80 citations
- DSE-GAN: Dynamic Semantic Evolution Generative Adversarial Network for Text-to-Image GenerationMengqi Huang, Zhendong Mao, Penghui Wang, Quan Wang et al.ACM MM 2022 · 26 citations
- Dual Attention GANs for Semantic Image SynthesisHao Tang, Song Bai, Nicu SebeACM MM 2020 · 81 citations
- ManiGAN: Text-Guided Image ManipulationBowen Li, Xiaojuan Qi, Thomas Lukasiewicz, Philip H. S. TorrCVPR 2020
- Adma-GAN: Attribute-Driven Memory Augmented GANs for Text-to-Image GenerationXintian Wu, Hanbin Zhao, Liangli Zheng, Shouhong Ding et al.ACM MM 2022 · 17 citations
