Simultaneous Machine Translation with Visual Context
Ozan Caglayan, Julia Ive, Veneta Haralampieva, Pranava Madhyastha, Loïc Barrault, Lucia Specia
Abstract
Simultaneous machine translation (SiMT) aims to translate a continuous input text stream into another language with the lowest latency and highest quality possible. The translation thus has to start with an incomplete source text, which is read progressively, creating the need for anticipation. In this paper, we seek to understand whether the addition of visual information can compensate for the missing source context. To this end, we analyse the impact of different multimodal approaches and visual features on state-of-the-art SiMT frameworks. Our results show that visual context is helpful and that visually-grounded models based on explicit object region information are much better than commonly used global features, reaching up to 3 BLEU points improvement under low latency scenarios. Our qualitative analysis illustrates cases where only the multimodal systems are able to translate correctly from English into gender-marked languages, as well as deal with differences in word order, such as adjective-noun placement between English and French.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext eb21612a-76d3-4e48-8ca8-3344819fc86dCited by top-tier papers7
- A Generative Framework for Simultaneous Machine TranslationYishu Miao, Phil Blunsom, Lucia SpeciaEMNLP 2021 · 12 citations
- Exploring Better Text Image Translation with Multimodal CodebookZhibin Lan, Jiawei Yu, Xiang Li, Wen Zhang et al.ACL 2023 · 12 citations
- SimQA: Detecting Simultaneous MT Errors through Word-by-Word Question AnsweringHyoJung Han, Marine Carpuat, Jordan L. Boyd-GraberEMNLP 2022 · 3 citations
- Toward Machine Interpreting: Lessons from Human Interpreting StudiesMatthias Sperber, Maureen de Seyssel, Jiajun Bao, Matthias PaulikEMNLP 2025 · 1 citation
- BERTGen: Multi-task Generation through BERTFaidon Mitzalis, Ozan Caglayan, Pranava Madhyastha, Lucia SpeciaACL 2021
Related papers
- Soul-Mix: Enhancing Multimodal Machine Translation with Manifold MixupXuxin Cheng, Ziyu Yao, Yifei Xin, Hao An et al.ACL 2024 · 3 citations
- Universal Simultaneous Machine Translation with Mixture-of-Experts Wait-k PolicyShaolei Zhang, Yang FengEMNLP 2021 · 19 citations
- Reducing Position Bias in Simultaneous Machine Translation with Length-Aware FrameworkShaolei Zhang, Yang FengACL 2022 · 23 citations
- SimulSLT: End-to-End Simultaneous Sign Language TranslationAoxiong Yin, Zhou Zhao, Jinglin Liu, Weike Jin et al.ACM MM 2021 · 35 citations
- Context Consistency between Training and Inference in Simultaneous Machine TranslationMeizhi Zhong, Lemao Liu, Kehai Chen, Mingming Yang et al.ACL 2024
