Understanding the Properties of Minimum Bayes Risk Decoding in Neural Machine Translation
Mathias Müller, Rico Sennrich
Abstract
Neural Machine Translation (NMT) currently exhibits biases such as producing translations that are too short and overgenerating frequent words, and shows poor robustness to copy noise in training data or domain shift. Recent work has tied these shortcomings to beam search -the de facto standard inference algorithm in NMT -and Eikema and Aziz (2020) propose to use Minimum Bayes Risk (MBR) decoding on unbiased samples instead. In this paper, we empirically investigate the properties of MBR decoding on a number of previously reported biases and failure cases of beam search. We find that MBR still exhibits a length and token frequency bias, owing to the MT metrics used as utility functions, but that MBR also increases robustness against copy noise in the training data and domain shift. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4484921e-a2f9-41a2-a3c9-8d468aea1f64Cited by top-tier papers18
- Detecting and Mitigating Hallucinations in Machine Translation: Model Internal Workings Alone Do Well, Sentence Similarity Even BetterDavid Dale, Elena Voita, Loïc Barrault, Marta R. Costa-jussàACL 2023 · 25 citations
- Efficient Minimum Bayes Risk Decoding using Low-Rank Matrix Completion AlgorithmsFiras Trabelsi, David Vilar, Mara Finkelstein, Markus FreitagNeurIPS 2024 · 18 citations
- Uncertainty Determines the Adequacy of the Mode and the Tractability of Decoding in Sequence-to-Sequence ModelsFelix Stahlberg, Ilia Kulikov, Shankar KumarACL 2022 · 13 citations
- Sampling-Based Approximations to Minimum Bayes Risk Decoding for Neural Machine TranslationBryan Eikema, Wilker AzizEMNLP 2022 · 10 citations
- Model-Based Minimum Bayes Risk Decoding for Text GenerationYuu Jinnai, Tetsuro Morimura, Ukyo Honda, Kaito Ariu et al.ICML 2024 · 9 citations
Builds on3
- ParaCrawl: Web-Scale Acquisition of Parallel CorporaMarta Bañón, Pinzhen Chen, Barry Haddow, Kenneth Heafield et al.ACL 2020 · 132 citations
- SSMBA: Self-Supervised Manifold Based Data Augmentation for Improving Out-of-Domain RobustnessNathan Ng, Kyunghyun Cho, Marzyeh GhassemiEMNLP 2020 · 6 citations
- COMET: A Neural Framework for MT EvaluationRicardo Rei, Craig Stewart, Ana C. Farinha, Alon LavieEMNLP 2020 · 6 citations
Related papers
- Unveiling the Power of Source: Source-based Minimum Bayes Risk Decoding for Neural Machine TranslationBoxuan Lyu, Hidetaka Kamigaito, Kotaro Funakoshi, Manabu OkumuraACL 2025
- Quality-Aware Translation Models: Efficient Generation and Quality Estimation in a Single ModelChristian Tomani, David Vilar, Markus Freitag, Colin Cherry et al.ACL 2024
- Digging Errors in NMT: Evaluating and Understanding Model Errors from Partial Hypothesis SpaceJianhao Yan, Chenming Wu, Fandong Meng, Jie ZhouEMNLP 2022
- Multi-Sentence Resampling: A Simple Approach to Alleviate Dataset Length Bias and Beam-Search DegradationIvan Provilkov, Andrey MalininEMNLP 2021 · 3 citations
- MBR and QE Finetuning: Training-time Distillation of the Best and Most Expensive Decoding MethodsMara Finkelstein, Markus FreitagICLR 2024 · 39 citations
