Dynamic Context-guided Capsule Network for Multimodal Machine Translation
Huan Lin, Fandong Meng, Jinsong Su, Yongjing Yin, Zhengyuan Yang, Yubin Ge, Jie Zhou, Jiebo Luo
Abstract
Multimodal machine translation (MMT), which mainly focuses on enhancing text-only translation with visual features, has attracted considerable attention from both computer vision and natural language processing communities. Most current MMT models resort to attention mechanism, global context modeling or multimodal joint representation learning to utilize visual features. However, the attention mechanism lacks sufficient semantic interactions between modalities while the other two provide fixed visual context, which is unsuitable for modeling the observed variability when generating translation. To address the above issues, in this paper, we propose a novel Dynamic Context-guided Capsule Network (DCCN) for MMT. Specifically, at each timestep of decoding, we first employ the conventional source-target attention to produce a timestep-specific source-side context vector. Next, DCCN takes this vector as input and uses it to guide the iterative extraction of related visual features via a context-guided dynamic routing mechanism. Particularly, we represent the input image with global and regional visual features, we introduce two parallel DCCNs to model multimodal context vectors with visual features at different granularities. Finally, we obtain two multimodal context vectors, which are fused and incorporated into the decoder for the prediction of the target word. Experimental results on the Multi30K dataset of English-to-German and English-to-French translation demonstrate the superiority of DCCN. Our code is available on https://github.com/DeepLearnXMU/MM-DCCN.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5d931bf8-6df2-480c-a069-cd4083043eedCited by top-tier papers21
- On Vision Features in Multimodal Machine TranslationBei Li, Chuanhao Lv, Zefan Zhou, Tao Zhou et al.ACL 2022 · 82 citations
- Multi-Modal Latent Space Learning for Chain-of-Thought Reasoning in Language ModelsLiqi He, Zuchao Li, Xiantao Cai, Ping WangAAAI 2024 · 38 citations
- VALHALLA: Visual Hallucination for Machine TranslationYi Li, Rameswar Panda, Yoon Kim, Chun-Fu Richard Chen et al.CVPR 2022 · 31 citations
- CLIPTrans: Transferring Visual Knowledge with Pre-trained Models for Multimodal Machine TranslationDevaansh Gupta, Siddhant Kharbanda, Jiawei Zhou, Wanhua Li et al.ICCV 2023 · 28 citations
- MSCTD: A Multimodal Sentiment Chat Translation DatasetYunlong Liang, Fandong Meng, Jinan Xu, Yufeng Chen et al.ACL 2022 · 27 citations
Builds on2
- A Novel Graph-based Multi-modal Fusion Encoder for Neural Machine TranslationYongjing Yin, Fandong Meng, Jinsong Su, Chulun Zhou et al.ACL 2020 · 145 citations
- Dual Directed Capsule Network for Very Low Resolution Image RecognitionManeet Singh, Shruti Nagpal, Richa Singh, Mayank VatsaICCV 2019 · 56 citations
Related papers
- Neural Machine Translation with Phrase-Level Universal Visual RepresentationsQingkai Fang, Yang FengACL 2022
- Efficient Object-Level Visual Context Modeling for Multimodal Machine Translation: Masking Irrelevant Objects Helps GroundingDexin Wang, Deyi XiongAAAI 2021 · 45 citations
- Visual Agreement Regularized Training for Multi-Modal Machine TranslationPengcheng Yang, Boxing Chen, Pei Zhang, Xu SunAAAI 2020 · 34 citations
- Leveraging Weighted Cross-Graph Attention for Visual and Semantic Enhanced Video Captioning NetworkDeepali Verma, Arya Haldar, Tanima DuttaAAAI 2023 · 13 citations
- Virtual Visual-Guided Domain-Shadow Fusion via Modal Exchanging for Domain-Specific Multi-Modal Neural Machine TranslationZhenyu Hou, Junjun GuoACM MM 2024 · 4 citations
