DSE-GAN: Dynamic Semantic Evolution Generative Adversarial Network for Text-to-Image Generation
Mengqi Huang, Zhendong Mao, Penghui Wang, Quan Wang, Yongdong Zhang
摘要
Text-to-image generation aims at generating realistic images which are semantically consistent with the given text. Previous works mainly adopt the multi-stage architecture by stacking generator-discriminator pairs to engage multiple adversarial training, where the text semantics used to provide generation guidance remain static across all stages. This work argues that text features at each stage should be adaptively re-composed conditioned on the status of the historical stage (.e., historical stage's text and image features) to provide diversified and accurate semantic guidance during the coarse-to-fine generation process. We thereby propose a novel Dynamical Semantic Evolution GAN (DSE-GAN) to re-compose each stage's text features under a novel single adversarial multi-stage architecture. Specifically, we design (1) Dynamic Semantic Evolution (DSE) module, which first aggregates historical image features to summarize the generative feedback, and then dynamically selects words required to be re-composed at each stage as well as re-composed them by dynamically enhancing or suppressing different granularity subspace's semantics. (2) Single Adversarial Multi-stage Architecture (SAMA), which extends the previous structure by eliminating complicated multiple adversarial training requirements and therefore allows more stages of text-image interactions, and finally facilitates the DSE module. We conduct comprehensive experiments and show that DSE-GAN achieves 7.48% and 37.8% relative FID improvement on two widely used benchmarks, i.e., CUB-200 and MSCOCO, respectively.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- DreamIdentity: Enhanced Editability for Efficient Face-Identity Preserved Image GenerationZhuowei Chen, Shancheng Fang, Wei Liu, Qian He 等AAAI 2024 · 被引用 26 次
- CustomContrast: A Multilevel Contrastive Perspective for Subject-Driven Text-to-Image CustomizationNan Chen, Mengqi Huang, Zhuowei Chen, Yang Zheng 等AAAI 2025 · 被引用 9 次
- Bridging the Modality Gap: Dimension Information Alignment and Sparse Spatial Constraint for Image-Text MatchingXiang Ma, Xuemei Li, Lexin Fang, Caiming ZhangACM MM 2024 · 被引用 4 次
- DualReal: Adaptive Joint Training for Lossless Identity-Motion Fusion in Video CustomizationWenchuan Wang, Mengqi Huang, Yijing Tu, Zhendong MaoICCV 2025 · 被引用 3 次
- When Measures are Unreliable: Imperceptible Adversarial Perturbations toward Top-k Multi-Label LearningYuchen Sun, Qianqian Xu, Zitai Wang, Qingming HuangACM MM 2023 · 被引用 2 次
它引用的顶会 Paper21
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray 等ICML 2021 · 被引用 6,356 次
- On Layer Normalization in the Transformer ArchitectureRuibin Xiong, Yunchang Yang, Di He, Kai Zheng 等ICML 2020 · 被引用 1,388 次
- CogView: Mastering Text-to-Image Generation via TransformersMing Ding, Zhuoyi Yang, Wenyi Hong, Wendi Zheng 等NeurIPS 2021 · 被引用 1,026 次
- TransGAN: Two Pure Transformers Can Make One Strong GAN, and That Can Scale UpYifan Jiang, Shiyu Chang, Zhangyang WangNeurIPS 2021 · 被引用 515 次
相关 Paper
- Semantics-Enhanced Adversarial Nets for Text-to-Image SynthesisHongchen Tan, Xiuping Liu, Xin Li, Yi Zhang 等ICCV 2019 · 被引用 80 次
- DF-GAN: A Simple and Effective Baseline for Text-to-Image SynthesisMing Tao, Hao Tang, Fei Wu, Xiaoyuan Jing 等CVPR 2022 · 被引用 296 次
- DAE-GAN: Dynamic Aspect-aware GAN for Text-to-Image SynthesisShulan Ruan, Yong Zhang, Kun Zhang, Yanbo Fan 等ICCV 2021 · 被引用 125 次
- R-GAN: Exploring Human-like Way for Reasonable Text-to-Image Synthesis via Generative Adversarial NetworksYanyuan Qiao, Qi Chen, Chaorui Deng, Ning Ding 等ACM MM 2021 · 被引用 18 次
- Cycle-Consistent Inverse GAN for Text-to-Image SynthesisHao Wang, Guosheng Lin, Steven C. H. Hoi, Chunyan MiaoACM MM 2021 · 被引用 47 次
