When Does Translation Require Context? A Data-driven, Multilingual Exploration
Patrick Fernandes, Kayo Yin, Emmy Liu, André F. T. Martins, Graham Neubig
Abstract
Although proper handling of discourse significantly contributes to the quality of machine translation (MT), these improvements are not adequately measured in common translation quality metrics. Recent works in context-aware MT attempt to target a small set of discourse phenomena during evaluation, however not in a fully systematic way. In this paper, we develop the Multilingual Discourse-Aware (MUDA) benchmark, a series of taggers that identify and evaluate model performance on discourse phenomena in any given dataset. The choice of phenomena is inspired by a novel methodology to systematically identify translations requiring context. We confirm the difficulty of previously studied phenomena while uncovering others that were previously unaddressed. We find that common context-aware MT models make only marginal improvements over context-agnostic models, which suggests these models do not handle these ambiguities effectively. We release code and data for 14 language pairs to encourage the MT community to focus on accurately capturing discourse phenomena. 1 * Equal contribution 1 Code available at https://github.com/CoderPat/MuDA . See §A for example usages of our released toolkit
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c68b75f7-5edb-407d-91c6-5bb795b69ff1Cited by top-tier papers7
- Paraphrasing evades detectors of AI-generated text, but retrieval is an effective defenseKalpesh Krishna, Yixiao Song, Marzena Karpinska, John Wieting et al.NeurIPS 2023 · 657 citations
- Quantifying the Plausibility of Context Reliance in Neural Machine TranslationGabriele Sarti, Grzegorz Chrupala, Malvina Nissim, Arianna BisazzaICLR 2024 · 8 citations
- Reconsidering Sentence-Level Sign Language TranslationGarrett Tanzer, Maximus Shengelia, Ken Harrenstien, David UthusEMNLP 2024 · 4 citations
- What Is That Talk About? A Video-to-Text Summarization Dataset for Scientific PresentationsDongqi Liu, Chenxi Whitehouse, Xi Yu, Louis Mahon et al.ACL 2025
- You Are What You Train: Effects of Data Composition on Training Context-aware Machine Translation ModelsPawel Maka, Yusuf Can Semerci, Jan Scholtes, Gerasimos SpanakisEMNLP 2025
Builds on3
- COMET: A Neural Framework for MT EvaluationRicardo Rei, Craig Stewart, Ana C. Farinha, Alon LavieEMNLP 2020 · 6 citations
- Measuring and Increasing Context Usage in Context-Aware Machine TranslationPatrick Fernandes, Kayo Yin, Graham Neubig, André F. T. MartinsACL 2021
- Do Context-Aware Translation Models Pay the Right Attention?Kayo Yin, Patrick Fernandes, Danish Pruthi, Aditi Chaudhary et al.ACL 2021
Related papers
- Discourse-Centric Evaluation of Document-level Machine Translation with a New Densely Annotated Parallel Corpus of NovelsYuchen Eleanor Jiang, Tianyu Liu, Shuming Ma, Dongdong Zhang et al.ACL 2023 · 7 citations
- DiscoX: Benchmarking Discourse-Level Translation in Expert DomainsXiying ZHAO, Zhoufutu Wen, Zhixuan Chen, Jingzhe Ding et al.ICLR 2026 · 2 citations
- Document-Level Machine Translation with Large-Scale Public Parallel CorporaProyag Pal, Alexandra Birch, Kenneth HeafieldACL 2024
- Towards Fully Automated Manga TranslationRyota Hinami, Shonosuke Ishiwatari, Kazuhiko Yasuda, Yusuke MatsuiAAAI 2021 · 41 citations
- Extrinsic Evaluation of Machine Translation MetricsNikita Moghe, Tom Sherborne, Mark Steedman, Alexandra BirchACL 2023 · 12 citations
