Increasing Visual Awareness in Multimodal Neural Machine Translation from an Information Theoretic Perspective
Baijun Ji, Tong Zhang, Yicheng Zou, Bojie Hu, Si Shen
摘要
Multimodal machine translation (MMT) aims to improve translation quality by equipping the source sentence with its corresponding image. Despite the promising performance, MMT models still suffer the problem of input degradation: models focus more on textual information while visual information is generally overlooked. In this paper, we endeavor to improve MMT performance by increasing visual awareness from an information theoretic perspective. In detail, we decompose the informative visual signals into two parts: source-specific information and target-specific information. We use mutual information to quantify them and propose two methods for objective optimization to better leverage visual signals. Experiments on two datasets demonstrate that our approach can effectively enhance the visual awareness of MMT model and achieve superior results against strong baselines.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- InfoPrompt: Information-Theoretic Soft Prompt Tuning for Natural Language UnderstandingJunda Wu, Tong Yu, Rui Wang, Zhao Song 等NeurIPS 2023 · 被引用 48 次
- Balancing Multimodal Training Through Game-Theoretic RegularizationKonstantinos Kontras, Thomas Strypsteen, Christos Chatzichristos, Paul Pu Liang 等NeurIPS 2025 · 被引用 17 次
- Bridging the Gap between Synthetic and Authentic Images for Multimodal Machine TranslationWenyu Guo, Qingkai Fang, Dong Yu, Yang FengEMNLP 2023 · 被引用 5 次
- Multimodal Neural Machine Translation: A Survey of the State of the ArtYi Feng, Chuanyi Li, Jiatong He, Zhenyu Hou 等EMNLP 2025 · 被引用 1 次
- Conditional Information Bottleneck for Multimodal Fusion: Overcoming Shortcut Learning in Sarcasm DetectionYihua Wang, Qi Jia, Cong Xu, Feiyu Chen 等AAAI 2026
它引用的顶会 Paper13
- A Mutual Information Maximization Perspective of Language Representation LearningLingpeng Kong, Cyprien de Masson d'Autume, Lei Yu, Wang Ling 等ICLR 2020 · 被引用 179 次
- A Novel Graph-based Multi-modal Fusion Encoder for Neural Machine TranslationYongjing Yin, Fandong Meng, Jinsong Su, Chulun Zhou 等ACL 2020 · 被引用 145 次
- Neural Machine Translation with Universal Visual RepresentationZhuosheng Zhang, Kehai Chen, Rui Wang, Masao Utiyama 等ICLR 2020 · 被引用 117 次
- On Vision Features in Multimodal Machine TranslationBei Li, Chuanhao Lv, Zefan Zhou, Tao Zhou 等ACL 2022 · 被引用 82 次
- Dynamic Context-guided Capsule Network for Multimodal Machine TranslationHuan Lin, Fandong Meng, Jinsong Su, Yongjing Yin 等ACM MM 2020 · 被引用 57 次
相关 Paper
- Neural Machine Translation with Phrase-Level Universal Visual RepresentationsQingkai Fang, Yang FengACL 2022
- Visual Agreement Regularized Training for Multi-Modal Machine TranslationPengcheng Yang, Boxing Chen, Pei Zhang, Xu SunAAAI 2020 · 被引用 34 次
- Soul-Mix: Enhancing Multimodal Machine Translation with Manifold MixupXuxin Cheng, Ziyu Yao, Yifei Xin, Hao An 等ACL 2024 · 被引用 3 次
- Good for Misconceived Reasons: An Empirical Revisiting on the Need for Visual Context in Multimodal Machine TranslationZhiyong Wu, Lingpeng Kong, Wei Bi, Xiang Li 等ACL 2021
- Efficient Object-Level Visual Context Modeling for Multimodal Machine Translation: Masking Irrelevant Objects Helps GroundingDexin Wang, Deyi XiongAAAI 2021 · 被引用 45 次
