On Vision Features in Multimodal Machine Translation
Bei Li, Chuanhao Lv, Zefan Zhou, Tao Zhou, Tong Xiao, Anxiang Ma, Jingbo Zhu
Abstract
Previous work on multimodal machine translation (MMT) has focused on the way of incorporating vision features into translation but little attention is on the quality of vision models. In this work, we investigate the impact of vision models on MMT. Given the fact that Transformer is becoming popular in computer vision, we experiment with various strong models (such as Vision Transformer) and enhanced features (such as object-detection and image captioning). We develop a selective attention model to study the patch-level contribution of an image in MMT. On detailed probing tasks, we find that stronger vision models are helpful for learning translation from the visual modality. Our results also suggest the need of carefully examining MMT models, especially when current benchmarks are small-scale and biased. Our code could be found at https: //github.com/libeineu/fairseq_mmt.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cd3d8248-df11-46bf-9e54-956b71253af5Cited by top-tier papers12
- KAM-CoT: Knowledge Augmented Multimodal Chain-of-Thoughts ReasoningDebjyoti Mondal, Suraj Modi, Subhadarshi Panda, Rituraj Singh et al.AAAI 2024 · 96 citations
- T-SciQ: Teaching Multimodal Chain-of-Thought Reasoning via Large Language Model Signals for Science Question AnsweringLei Wang, Yi Hu, Jiabang He, Xing Xu et al.AAAI 2024 · 95 citations
- Multi-Modal Latent Space Learning for Chain-of-Thought Reasoning in Language ModelsLiqi He, Zuchao Li, Xiantao Cai, Ping WangAAAI 2024 · 38 citations
- Exploring Better Text Image Translation with Multimodal CodebookZhibin Lan, Jiawei Yu, Xiang Li, Wen Zhang et al.ACL 2023 · 12 citations
- PEIT: Bridging the Modality Gap with Pre-trained Models for End-to-End Image TranslationShaolin Zhu, Shangjie Li, Yikun Lei, Deyi XiongACL 2023 · 12 citations
Builds on7
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- A Novel Graph-based Multi-modal Fusion Encoder for Neural Machine TranslationYongjing Yin, Fandong Meng, Jinsong Su, Chulun Zhou et al.ACL 2020 · 145 citations
- Neural Machine Translation with Universal Visual RepresentationZhuosheng Zhang, Kehai Chen, Rui Wang, Masao Utiyama et al.ICLR 2020 · 117 citations
- Dynamic Context-guided Capsule Network for Multimodal Machine TranslationHuan Lin, Fandong Meng, Jinsong Su, Yongjing Yin et al.ACM MM 2020 · 57 citations
Related papers
- Increasing Visual Awareness in Multimodal Neural Machine Translation from an Information Theoretic PerspectiveBaijun Ji, Tong Zhang, Yicheng Zou, Bojie Hu et al.EMNLP 2022 · 11 citations
- Efficient Object-Level Visual Context Modeling for Multimodal Machine Translation: Masking Irrelevant Objects Helps GroundingDexin Wang, Deyi XiongAAAI 2021 · 45 citations
- Good for Misconceived Reasons: An Empirical Revisiting on the Need for Visual Context in Multimodal Machine TranslationZhiyong Wu, Lingpeng Kong, Wei Bi, Xiang Li et al.ACL 2021
- LVP-M3: Language-aware Visual Prompt for Multilingual Multimodal Machine TranslationHongcheng Guo, Jiaheng Liu, Haoyang Huang, Jian Yang et al.EMNLP 2022 · 9 citations
- Neural Machine Translation with Phrase-Level Universal Visual RepresentationsQingkai Fang, Yang FengACL 2022
