DF-GAN: A Simple and Effective Baseline for Text-to-Image Synthesis
Ming Tao, Hao Tang, Fei Wu, Xiaoyuan Jing, Bing-Kun Bao, Changsheng Xu
摘要
Synthesizing high-quality realistic images from text descriptions is a challenging task. Existing text-to-image Generative Adversarial Networks generally employ a stacked architecture as the backbone yet still remain three flaws. First, the stacked architecture introduces the entanglements between generators of different image scales. Second, existing studies prefer to apply and fix extra networks in adversarial learning for text-image semantic consistency, which limits the supervision capability of these networks. Third, the cross-modal attention-based text-image fusion that widely adopted by previous works is limited on several special image scales because of the computational cost. To these ends, we propose a simpler but more effective Deep Fusion Generative Adversarial Networks (DF-GAN). To be specific, we propose: (i) a novel one-stage text-to-image backbone that directly synthesizes high-resolution images without entanglements between different generators, (ii) a novel Target-Aware Discriminator composed of Matching-Aware Gradient Penalty and One-Way Output, which enhances the text-image semantic consistency without introducing extra networks, (iii) a novel deep text-image fusion block, which deepens the fusion process to make a full fusion between text and visual features. Compared with current state-of-the-art methods, our proposed DF-GAN is simpler but more efficient to synthesize realistic and text-matching images and achieves better performance on widely used datasets. Code is available at https: //github.com/tobran/DF-GAN .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper89
- MasaCtrl: Tuning-Free Mutual Self-Attention Control for Consistent Image Synthesis and EditingMingdeng Cao, Xintao Wang, Zhongang Qi, Ying Shan 等ICCV 2023 · 被引用 770 次
- Muse: Text-To-Image Generation via Masked Generative TransformersHuiwen Chang, Han Zhang, Jarred Barber, Aaron Maschinot 等ICML 2023 · 被引用 751 次
- Attend-and-Excite: Attention-Based Semantic Guidance for Text-to-Image Diffusion ModelsHila Chefer, Yuval Alaluf, Yael Vinker, Lior Wolf 等SIGGRAPH 2023 · 被引用 438 次
- InstaFlow: One Step is Enough for High-Quality Diffusion-Based Text-to-Image GenerationXingchao Liu, Xiwen Zhang, Jianzhu Ma, Jian Peng 等ICLR 2024 · 被引用 358 次
- Optimizing Prompts for Text-to-Image GenerationYaru Hao, Zewen Chi, Li Dong, Furu WeiNeurIPS 2023 · 被引用 303 次
它引用的顶会 Paper8
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray 等ICML 2021 · 被引用 6,356 次
- CogView: Mastering Text-to-Image Generation via TransformersMing Ding, Zhuoyi Yang, Wenyi Hong, Wendi Zheng 等NeurIPS 2021 · 被引用 1,026 次
- DAE-GAN: Dynamic Aspect-aware GAN for Text-to-Image SynthesisShulan Ruan, Yong Zhang, Kun Zhang, Yanbo Fan 等ICCV 2021 · 被引用 125 次
- TIME: Text and Image Mutual-Translation Adversarial NetworksBingchen Liu, Kunpeng Song, Yizhe Zhu, Gerard de Melo 等AAAI 2021 · 被引用 35 次
相关 Paper
- Semantics-Enhanced Adversarial Nets for Text-to-Image SynthesisHongchen Tan, Xiuping Liu, Xin Li, Yi Zhang 等ICCV 2019 · 被引用 80 次
- DSE-GAN: Dynamic Semantic Evolution Generative Adversarial Network for Text-to-Image GenerationMengqi Huang, Zhendong Mao, Penghui Wang, Quan Wang 等ACM MM 2022 · 被引用 26 次
- RiFeGAN: Rich Feature Generation for Text-to-Image Synthesis From Prior KnowledgeJun Cheng, Fuxiang Wu, Yanling Tian, Lei Wang 等CVPR 2020
- Rethinking Super-Resolution as Text-Guided Details GenerationChenxi Ma, Bo Yan, Qing Lin, Weimin Tan 等ACM MM 2022 · 被引用 6 次
- SeD: Semantic-Aware Discriminator for Image Super-ResolutionBingchen Li, Xin Li, Hanxin Zhu, Yeying Jin 等CVPR 2024
