TIME: Text and Image Mutual-Translation Adversarial Networks
Bingchen Liu, Kunpeng Song, Yizhe Zhu, Gerard de Melo, Ahmed Elgammal
Abstract
Focusing on text-to-image (T2I) generation, we propose Text and Image Mutual-Translation Adversarial Networks (TIME), a lightweight but effective model that jointly learns a T2I generator G and an image captioning discriminator D under the Generative Adversarial Network framework. While previous methods tackle the T2I problem as a uni-directional task and use pre-trained language models to enforce the image-text consistency, TIME requires neither extra modules nor pretraining. We show that the performance of G can be boosted substantially by training it jointly with D as a language model. Specifically, we adopt Transformers to model the cross-modal connections between the image features and word embeddings, and design an annealing conditional hinge loss that dynamically balances the adversarial learning. In our experiments, TIME achieves state-of-the-art (SOTA) performance on the CUB and MS-COCO dataset (Inception Score of 4.91 and Fréchet Inception Distance of 14.3 on CUB), and shows promising performance on MS-COCO on image captioning and downstream vision-language tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers7
- Towards Faster and Stabilized GAN Training for High-fidelity Few-shot Image SynthesisBingchen Liu, Yizhe Zhu, Kunpeng Song, Ahmed ElgammalICLR 2021 · 307 citations
- DF-GAN: A Simple and Effective Baseline for Text-to-Image SynthesisMing Tao, Hao Tang, Fei Wu, Xiaoyuan Jing et al.CVPR 2022 · 296 citations
- L-CoDe: Language-Based Colorization Using Color-Object Decoupled ConditionsShuchen Weng, Hao Wu, Zheng Chang, Jiajun Tang et al.AAAI 2022 · 57 citations
- DSE-GAN: Dynamic Semantic Evolution Generative Adversarial Network for Text-to-Image GenerationMengqi Huang, Zhendong Mao, Penghui Wang, Quan Wang et al.ACM MM 2022 · 26 citations
- Fonts Like This but Happier: A New Way to Discover FontsTugba Kulahcioglu, Gerard de MeloACM MM 2020 · 14 citations
Related papers
- Translation-Enhanced Multilingual Text-to-Image GenerationYaoyiran Li, Ching-Yun Chang, Stephen Rawls, Ivan Vulic et al.ACL 2023 · 8 citations
- E2E-VLP: End-to-End Vision-Language Pre-training Enhanced by Visual LearningHaiyang Xu, Ming Yan, Chenliang Li, Bin Bi et al.ACL 2021
- Adma-GAN: Attribute-Driven Memory Augmented GANs for Text-to-Image GenerationXintian Wu, Hanbin Zhao, Liangli Zheng, Shouhong Ding et al.ACM MM 2022 · 17 citations
- Unified Discrete Diffusion for Simultaneous Vision-Language GenerationMinghui Hu, Chuanxia Zheng, Zuopeng Yang, Tat-Jen Cham et al.ICLR 2023 · 8 citations
- Unified Vision-Language Pre-Training for Image Captioning and VQALuowei Zhou, Hamid Palangi, Lei Zhang, Houdong Hu et al.AAAI 2020 · 1,047 citations
