Translation-Enhanced Multilingual Text-to-Image Generation
Yaoyiran Li, Ching-Yun Chang, Stephen Rawls, Ivan Vulic, Anna Korhonen
摘要
Research on text-to-image generation (TTI) still predominantly focuses on the English language due to the lack of annotated imagecaption data in other languages; in the long run, this might widen inequitable access to TTI technology. In this work, we thus investigate multilingual TTI (termed mTTI) and the current potential of neural machine translation (NMT) to bootstrap mTTI systems. We provide two key contributions. 1) Relying on a multilingual multi-modal encoder, we provide a systematic empirical study of standard methods used in cross-lingual NLP when applied to mTTI: TRANSLATE TRAIN, TRANS-LATE TEST, and ZERO-SHOT TRANSFER. 2) We propose Ensemble Adapter (ENSAD), a novel parameter-efficient approach that learns to weigh and consolidate the multilingual text knowledge within the mTTI framework, mitigating the language gap and thus improving mTTI performance. Our evaluations on standard mTTI datasets COCO-CN, Multi30K Task2, and LAION-5B demonstrate the potential of translation-enhanced mTTI systems and also validate the benefits of the proposed EN-SAD which derives consistent gains across all datasets. Further investigations on model variants, ablation studies, and qualitative analyses provide additional insights on the inner workings of the proposed mTTI approaches.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Where Culture Fades: Revealing the Cultural Gap in Text-to-Image GenerationChuancheng Shi, Shangze Li, Shiming Guo, Simiao Xie 等CVPR 2026 · 被引用 14 次
- Bridge Diffusion Model: Bridge Chinese Text-to-Image Diffusion Model with English CommunitiesShanyuan Liu, Bo Cheng, Yuhang Ma, Liebucha Wu 等AAAI 2025 · 被引用 5 次
- On Bilingual Lexicon Induction with Large Language ModelsYaoyiran Li, Anna Korhonen, Ivan VulicEMNLP 2023 · 被引用 2 次
它引用的顶会 Paper19
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray 等ICML 2021 · 被引用 6,356 次
- GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion ModelsAlexander Quinn Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam 等ICML 2022 · 被引用 4,691 次
- Training Generative Adversarial Networks with Limited DataTero Karras, Miika Aittala, Janne Hellsten, Samuli Laine 等NeurIPS 2020 · 被引用 2,345 次
相关 Paper
- TIME: Text and Image Mutual-Translation Adversarial NetworksBingchen Liu, Kunpeng Song, Yizhe Zhu, Gerard de Melo 等AAAI 2021 · 被引用 35 次
- MuLan: Adapting Multilingual Diffusion Models for Hundreds of Languages with Negligible CostSen Xing, Muyan Zhong, Zeqiang Lai, Liangchen Li 等ICML 2025
- Scaling Laws for Conditional Emergence of Multilingual Image Captioning via Generalization from TranslationJulian Spravil, Sebastian Houben, Sven BehnkeAAAI 2026
- LVP-M3: Language-aware Visual Prompt for Multilingual Multimodal Machine TranslationHongcheng Guo, Jiaheng Liu, Haoyang Huang, Jian Yang 等EMNLP 2022 · 被引用 9 次
- Learning Language Specific Sub-network for Multilingual Machine TranslationZehui Lin, Liwei Wu, Mingxuan Wang, Lei LiACL 2021
