Make Imagination Clearer! Stable Diffusion-based Visual Imagination for Multimodal Machine Translation
Andong Chen, Yuchen Song, Kehai Chen, Xuefeng Bai, Muyun Yang, Liqiang Nie, Jie Liu, Tiejun Zhao, Min Zhang
Abstract
Visual information has been introduced for enhancing machine translation (MT), and its effectiveness heavily relies on the availability of large amounts of bilingual parallel sentence pairs with manual image annotations. In this paper, we propose a stable diffusionbased imagination network integrated into a multimodal large language model (MLLM) to explicitly generate an image for each source sentence, thereby advancing multimodal MT. Particularly, we build heuristic feedback with reinforcement learning to ensure the consistency of the generated image with the source sentence without the supervision of visual information, which breaks the highcost bottleneck of image annotation in MT. Furthermore, the proposed method enables imaginative visual information to be integrated into text-only MT in addition to multimodal MT. Experimental results show that our model significantly outperforms existing multimodal MT and text-only MT, especially achieving an average improvement of more than 12 BLEU points on Multi30K and MSCOCO multimodal MT benchmarks. 1 * Corresponding author. 1 Our code is available at https://github.com/ coder109/IMAGE Three women, walking or standing near a wall, outside.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6d5586ec-7a16-48e5-a29f-9a7567c121bfCited by top-tier papers3
- PRIM: Towards Practical In-Image Multilingual Machine TranslationYanzhi Tian, Zeming Liu, Zhengyang Liu, Chong Feng et al.EMNLP 2025 · 4 citations
- Culture In a Frame: C3B as a Comic-Based Benchmark for Multimodal Culturally AwarenessYuchen Song, Andong Chen, Wenxin Zhu, Kehai Chen et al.ICLR 2026 · 3 citations
- Scalable Multilingual Multimodal Machine Translation with Speech-Text FusionYexing Du, Youcheng Pan, Zekun Wang, Zheng Chu et al.ICLR 2026
Builds on30
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
Related papers
- Rich Human Feedback for Text-to-Image GenerationYouwei Liang, Junfeng He, Gang Li, Peizhao Li et al.CVPR 2024
- VALHALLA: Visual Hallucination for Machine TranslationYi Li, Rameswar Panda, Yoon Kim, Chun-Fu Richard Chen et al.CVPR 2022 · 31 citations
- Neural Machine Translation with Phrase-Level Universal Visual RepresentationsQingkai Fang, Yang FengACL 2022
- Soul-Mix: Enhancing Multimodal Machine Translation with Manifold MixupXuxin Cheng, Ziyu Yao, Yifei Xin, Hao An et al.ACL 2024 · 3 citations
- Generative Multimodal Data Augmentation for Low-Resource Multimodal Named Entity RecognitionZiyan Li, Jianfei Yu, Jia Yang, Wenya Wang et al.ACM MM 2024 · 13 citations
