Increasing Visual Awareness in Multimodal Neural Machine Translation from an Information Theoretic Perspective
Baijun Ji, Tong Zhang, Yicheng Zou, Bojie Hu, Si Shen
Abstract
Multimodal machine translation (MMT) aims to improve translation quality by equipping the source sentence with its corresponding image. Despite the promising performance, MMT models still suffer the problem of input degradation: models focus more on textual information while visual information is generally overlooked. In this paper, we endeavor to improve MMT performance by increasing visual awareness from an information theoretic perspective. In detail, we decompose the informative visual signals into two parts: source-specific information and target-specific information. We use mutual information to quantify them and propose two methods for objective optimization to better leverage visual signals. Experiments on two datasets demonstrate that our approach can effectively enhance the visual awareness of MMT model and achieve superior results against strong baselines.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1ef6237b-b33a-445e-8ea7-bdb3ac08a23dCited by top-tier papers5
- InfoPrompt: Information-Theoretic Soft Prompt Tuning for Natural Language UnderstandingJunda Wu, Tong Yu, Rui Wang, Zhao Song et al.NeurIPS 2023 · 48 citations
- Balancing Multimodal Training Through Game-Theoretic RegularizationKonstantinos Kontras, Thomas Strypsteen, Christos Chatzichristos, Paul Pu Liang et al.NeurIPS 2025 · 17 citations
- Bridging the Gap between Synthetic and Authentic Images for Multimodal Machine TranslationWenyu Guo, Qingkai Fang, Dong Yu, Yang FengEMNLP 2023 · 5 citations
- Multimodal Neural Machine Translation: A Survey of the State of the ArtYi Feng, Chuanyi Li, Jiatong He, Zhenyu Hou et al.EMNLP 2025 · 1 citation
- Conditional Information Bottleneck for Multimodal Fusion: Overcoming Shortcut Learning in Sarcasm DetectionYihua Wang, Qi Jia, Cong Xu, Feiyu Chen et al.AAAI 2026
Builds on13
- A Mutual Information Maximization Perspective of Language Representation LearningLingpeng Kong, Cyprien de Masson d'Autume, Lei Yu, Wang Ling et al.ICLR 2020 · 179 citations
- A Novel Graph-based Multi-modal Fusion Encoder for Neural Machine TranslationYongjing Yin, Fandong Meng, Jinsong Su, Chulun Zhou et al.ACL 2020 · 145 citations
- Neural Machine Translation with Universal Visual RepresentationZhuosheng Zhang, Kehai Chen, Rui Wang, Masao Utiyama et al.ICLR 2020 · 117 citations
- On Vision Features in Multimodal Machine TranslationBei Li, Chuanhao Lv, Zefan Zhou, Tao Zhou et al.ACL 2022 · 82 citations
- Dynamic Context-guided Capsule Network for Multimodal Machine TranslationHuan Lin, Fandong Meng, Jinsong Su, Yongjing Yin et al.ACM MM 2020 · 57 citations
Related papers
- Neural Machine Translation with Phrase-Level Universal Visual RepresentationsQingkai Fang, Yang FengACL 2022
- Visual Agreement Regularized Training for Multi-Modal Machine TranslationPengcheng Yang, Boxing Chen, Pei Zhang, Xu SunAAAI 2020 · 34 citations
- Soul-Mix: Enhancing Multimodal Machine Translation with Manifold MixupXuxin Cheng, Ziyu Yao, Yifei Xin, Hao An et al.ACL 2024 · 3 citations
- Good for Misconceived Reasons: An Empirical Revisiting on the Need for Visual Context in Multimodal Machine TranslationZhiyong Wu, Lingpeng Kong, Wei Bi, Xiang Li et al.ACL 2021
- Efficient Object-Level Visual Context Modeling for Multimodal Machine Translation: Masking Irrelevant Objects Helps GroundingDexin Wang, Deyi XiongAAAI 2021 · 45 citations
