VALHALLA: Visual Hallucination for Machine Translation
Yi Li, Rameswar Panda, Yoon Kim, Chun-Fu Richard Chen, Rogério Feris, David D. Cox, Nuno Vasconcelos
Abstract
Designing better machine translation systems by considering auxiliary inputs such as images has attracted much attention in recent years. While existing methods show promising performance over the conventional text-only translation systems, they typically require paired text and image as input during inference, which limits their applicability to real-world scenarios. In this paper, we introduce a visual hallucination framework, called VALHALLA, which requires only source sentences at inference time and instead uses hallucinated visual representations for multi-modal machine translation. In particular, given a source sentence an autoregressive hallucination transformer is used to predict a discrete visual representation from the input text, and the combined text and hallucinated representations are utilized to obtain the target translation. We train the hallucination transformer jointly with the translation transformer using standard backpropagation with crossentropy losses while being guided by an additional loss that encourages consistency between predictions using either groundtruth or hallucinated visual representations. Extensive experiments on three standard translation datasets with a diverse set of language pairs demonstrate the effectiveness of our approach over both text-only baselines and state-of-the-art methods. Project page: http://www.svcl.ucsd.jects/valhalla.edu/pro.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 523d3bbf-2061-4790-8c8a-392f3e69147bCited by top-tier papers9
- Scene Graph as Pivoting: Inference-time Image-free Unsupervised Multimodal Machine Translation with Visual Scene HallucinationHao Fei, Qian Liu, Meishan Zhang, Min Zhang et al.ACL 2023 · 45 citations
- CLIPTrans: Transferring Visual Knowledge with Pre-trained Models for Multimodal Machine TranslationDevaansh Gupta, Siddhant Kharbanda, Jiawei Zhou, Wanhua Li et al.ICCV 2023 · 28 citations
- Tackling Ambiguity with Images: Improved Multimodal Machine Translation and Contrastive EvaluationMatthieu Futeral, Cordelia Schmid, Ivan Laptev, Benoît Sagot et al.ACL 2023 · 14 citations
- Reading Between the Heat: Co-Teaching Body Thermal Signatures for Non-intrusive Stress DetectionYi Xiao, Harshit Sharma, Zhongyang Zhang, Dessa Bergen-Cico et al.UbiComp 2024 · 10 citations
- Can visual language models resolve textual ambiguity with visual cues? Let visual puns tell you!Jiwan Chung, Seungwon Lim, Jaehyun Jeon, Seungbeen Lee et al.EMNLP 2024 · 8 citations
Builds on17
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
- VaTeX: A Large-Scale, High-Quality Multilingual Dataset for Video-and-Language ResearchXin Wang, Jiawei Wu, Jun-Kun Chen, Lei Li et al.ICCV 2019 · 688 citations
- Visual Semantic Reasoning for Image-Text MatchingKunpeng Li, Yulun Zhang, Kai Li, Yuanyuan Li et al.ICCV 2019 · 598 citations
- AdaShare: Learning What To Share For Efficient Deep Multi-Task LearningXimeng Sun, Rameswar Panda, Rogério Feris, Kate SaenkoNeurIPS 2020 · 337 citations
Related papers
- Neural Machine Translation with Phrase-Level Universal Visual RepresentationsQingkai Fang, Yang FengACL 2022
- Visual Agreement Regularized Training for Multi-Modal Machine TranslationPengcheng Yang, Boxing Chen, Pei Zhang, Xu SunAAAI 2020 · 34 citations
- Make Imagination Clearer! Stable Diffusion-based Visual Imagination for Multimodal Machine TranslationAndong Chen, Yuchen Song, Kehai Chen, Xuefeng Bai et al.ACL 2025
- Self-Correcting Decoding with Generative Feedback for Mitigating Hallucinations in Large Vision-Language ModelsCe Zhang, Zifu Wan, Zhehan Kan, Martin Q. Ma et al.ICLR 2025
- Bridging the Gap between Synthetic and Authentic Images for Multimodal Machine TranslationWenyu Guo, Qingkai Fang, Dong Yu, Yang FengEMNLP 2023 · 5 citations
