Efficient Object-Level Visual Context Modeling for Multimodal Machine Translation: Masking Irrelevant Objects Helps Grounding
Dexin Wang, Deyi Xiong
摘要
Visual context provides grounding information for multimodal machine translation (MMT). However, previous MMT models and probing studies on visual features suggest that visual information is less explored in MMT as it is often redundant to textual information. In this paper, we propose an Object-level Visual Context modeling framework (OVC) to efficiently capture and explore visual information for multimodal machine translation. With detected objects, the proposed OVC encourages MMT to ground translation on desirable visual objects by masking irrelevant objects in the visual modality. We equip the proposed with an additional object-masking loss to achieve this goal. The object-masking loss is estimated according to the similarity between masked objects and the source texts so as to encourage masking source-irrelevant objects. Additionally, in order to generate vision-consistent target words, we further propose a vision-weighted translation loss for OVC. Experiments on MMT datasets demonstrate that the proposed OVC model outperforms state-of-the-art MMT models and analyses show that masking irrelevant objects helps grounding in MMT.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- On Vision Features in Multimodal Machine TranslationBei Li, Chuanhao Lv, Zefan Zhou, Tao Zhou 等ACL 2022 · 被引用 82 次
- Distill The Image to Nowhere: Inversion Knowledge Distillation for Multimodal Machine TranslationRu Peng, Yawen Zeng, Jake ZhaoEMNLP 2022 · 被引用 15 次
- Increasing Visual Awareness in Multimodal Neural Machine Translation from an Information Theoretic PerspectiveBaijun Ji, Tong Zhang, Yicheng Zou, Bojie Hu 等EMNLP 2022 · 被引用 11 次
- Multimodal Neural Machine Translation: A Survey of the State of the ArtYi Feng, Chuanyi Li, Jiatong He, Zhenyu Hou 等EMNLP 2025 · 被引用 1 次
- SHIFT: Selected Helpful Informative Frame for Video-guided Machine TranslationBoyu Guan, Chuang Han, Yining Zhang, Yupu Liang 等EMNLP 2025
它引用的顶会 Paper4
- VaTeX: A Large-Scale, High-Quality Multilingual Dataset for Video-and-Language ResearchXin Wang, Jiawei Wu, Jun-Kun Chen, Lei Li 等ICCV 2019 · 被引用 688 次
- A Novel Graph-based Multi-modal Fusion Encoder for Neural Machine TranslationYongjing Yin, Fandong Meng, Jinsong Su, Chulun Zhou 等ACL 2020 · 被引用 145 次
- Neural Machine Translation with Universal Visual RepresentationZhuosheng Zhang, Kehai Chen, Rui Wang, Masao Utiyama 等ICLR 2020 · 被引用 117 次
- Visual Agreement Regularized Training for Multi-Modal Machine TranslationPengcheng Yang, Boxing Chen, Pei Zhang, Xu SunAAAI 2020 · 被引用 34 次
相关 Paper
- Neural Machine Translation with Phrase-Level Universal Visual RepresentationsQingkai Fang, Yang FengACL 2022
- UC2: Universal Cross-Lingual Cross-Modal Vision-and-Language Pre-TrainingMingyang Zhou, Luowei Zhou, Shuohang Wang, Yu Cheng 等CVPR 2021
- Improving Visual Grounding with Visual-Linguistic Verification and Iterative ReasoningLi Yang, Yan Xu, Chunfeng Yuan, Wei Liu 等CVPR 2022 · 被引用 146 次
- CL2CM: Improving Cross-Lingual Cross-Modal Retrieval via Cross-Lingual Knowledge TransferYabing Wang, Fan Wang, Jianfeng Dong, Hao LuoAAAI 2024 · 被引用 20 次
- Dynamic Context-guided Capsule Network for Multimodal Machine TranslationHuan Lin, Fandong Meng, Jinsong Su, Yongjing Yin 等ACM MM 2020 · 被引用 57 次
