Translation-Enhanced Multilingual Text-to-Image Generation
Yaoyiran Li, Ching-Yun Chang, Stephen Rawls, Ivan Vulic, Anna Korhonen
Abstract
Research on text-to-image generation (TTI) still predominantly focuses on the English language due to the lack of annotated imagecaption data in other languages; in the long run, this might widen inequitable access to TTI technology. In this work, we thus investigate multilingual TTI (termed mTTI) and the current potential of neural machine translation (NMT) to bootstrap mTTI systems. We provide two key contributions. 1) Relying on a multilingual multi-modal encoder, we provide a systematic empirical study of standard methods used in cross-lingual NLP when applied to mTTI: TRANSLATE TRAIN, TRANS-LATE TEST, and ZERO-SHOT TRANSFER. 2) We propose Ensemble Adapter (ENSAD), a novel parameter-efficient approach that learns to weigh and consolidate the multilingual text knowledge within the mTTI framework, mitigating the language gap and thus improving mTTI performance. Our evaluations on standard mTTI datasets COCO-CN, Multi30K Task2, and LAION-5B demonstrate the potential of translation-enhanced mTTI systems and also validate the benefits of the proposed EN-SAD which derives consistent gains across all datasets. Further investigations on model variants, ablation studies, and qualitative analyses provide additional insights on the inner workings of the proposed mTTI approaches.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c3370e6b-5861-4302-85dc-d8fee5add353Cited by top-tier papers3
- Where Culture Fades: Revealing the Cultural Gap in Text-to-Image GenerationChuancheng Shi, Shangze Li, Shiming Guo, Simiao Xie et al.CVPR 2026 · 14 citations
- Bridge Diffusion Model: Bridge Chinese Text-to-Image Diffusion Model with English CommunitiesShanyuan Liu, Bo Cheng, Yuhang Ma, Liebucha Wu et al.AAAI 2025 · 5 citations
- On Bilingual Lexicon Induction with Large Language ModelsYaoyiran Li, Anna Korhonen, Ivan VulicEMNLP 2023 · 2 citations
Builds on19
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
- GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion ModelsAlexander Quinn Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam et al.ICML 2022 · 4,691 citations
- Training Generative Adversarial Networks with Limited DataTero Karras, Miika Aittala, Janne Hellsten, Samuli Laine et al.NeurIPS 2020 · 2,345 citations
Related papers
- TIME: Text and Image Mutual-Translation Adversarial NetworksBingchen Liu, Kunpeng Song, Yizhe Zhu, Gerard de Melo et al.AAAI 2021 · 35 citations
- MuLan: Adapting Multilingual Diffusion Models for Hundreds of Languages with Negligible CostSen Xing, Muyan Zhong, Zeqiang Lai, Liangchen Li et al.ICML 2025
- Scaling Laws for Conditional Emergence of Multilingual Image Captioning via Generalization from TranslationJulian Spravil, Sebastian Houben, Sven BehnkeAAAI 2026
- LVP-M3: Language-aware Visual Prompt for Multilingual Multimodal Machine TranslationHongcheng Guo, Jiaheng Liu, Haoyang Huang, Jian Yang et al.EMNLP 2022 · 9 citations
- Learning Language Specific Sub-network for Multilingual Machine TranslationZehui Lin, Liwei Wu, Mingxuan Wang, Lei LiACL 2021
