Sampling-Based Approximations to Minimum Bayes Risk Decoding for Neural Machine Translation
Bryan Eikema, Wilker Aziz
Abstract
In NMT we search for the mode of the model distribution to form predictions. The mode and other high-probability translations found by beam search have been shown to often be inadequate in a number of ways. This prevents improving translation quality through better search, as these idiosyncratic translations end up selected by the decoding algorithm, a problem known as the beam search curse. Recently, an approximation to minimum Bayes risk (MBR) decoding has been proposed as an alternative decision rule that would likely not suffer from the same problems. We analyse this approximation and establish that it has no equivalent to the beam search curse. We then design approximations that decouple the cost of exploration from the cost of robust estimation of expected utility. This allows for much larger hypothesis spaces, which we show to be beneficial. We also show that mode-seeking strategies can aid in constructing compact sets of promising hypotheses and that MBR is effective in identifying good translations in them. We conduct experiments on three language pairs varying in amounts of resources available: English into and from German, Romanian, and Nepali. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1ea1e751-1194-434e-b9a1-0cea821efa17Cited by top-tier papers19
- MBR and QE Finetuning: Training-time Distillation of the Best and Most Expensive Decoding MethodsMara Finkelstein, Markus FreitagICLR 2024 · 39 citations
- Efficient Minimum Bayes Risk Decoding using Low-Rank Matrix Completion AlgorithmsFiras Trabelsi, David Vilar, Mara Finkelstein, Markus FreitagNeurIPS 2024 · 18 citations
- A Cheaper and Better Diffusion Language Model with Soft-Masked NoiseJiaao Chen, Aston Zhang, Mu Li, Alex Smola et al.EMNLP 2023 · 16 citations
- Model-Based Minimum Bayes Risk Decoding for Text GenerationYuu Jinnai, Tetsuro Morimura, Ukyo Honda, Kaito Ariu et al.ICML 2024 · 9 citations
- What Comes Next? Evaluating Uncertainty in Neural Text Generators Against Human Production VariabilityMario Giulianelli, Joris Baan, Wilker Aziz, Raquel Fernández et al.EMNLP 2023 · 5 citations
Builds on5
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- If beam search is the answer, what was the question?Clara Meister, Ryan Cotterell, Tim VieiraEMNLP 2020 · 26 citations
- Machine Translation Decoding beyond Beam SearchRémi Leblond, Jean-Baptiste Alayrac, Laurent Sifre, Miruna Pislar et al.EMNLP 2021 · 6 citations
- Energy-Based Reranking: Improving Neural Machine Translation Using Energy-Based ModelsSumanta Bhattacharyya, Amirmohammad Rooshenas, Subhajit Naskar, Simeng Sun et al.ACL 2021
- Understanding the Properties of Minimum Bayes Risk Decoding in Neural Machine TranslationMathias Müller, Rico SennrichACL 2021
Related papers
- Unveiling the Power of Source: Source-based Minimum Bayes Risk Decoding for Neural Machine TranslationBoxuan Lyu, Hidetaka Kamigaito, Kotaro Funakoshi, Manabu OkumuraACL 2025
- Theoretical Guarantees for Minimum Bayes Risk DecodingYuki Ichihara, Yuu Jinnai, Kaito Ariu, Tetsuro Morimura et al.ACL 2025
- Quality-Aware Translation Models: Efficient Generation and Quality Estimation in a Single ModelChristian Tomani, David Vilar, Markus Freitag, Colin Cherry et al.ACL 2024
- Case-Based Decision-Theoretic Decoding with Quality MemoriesHiroyuki Deguchi, Masaaki NagataEMNLP 2025
- Digging Errors in NMT: Evaluating and Understanding Model Errors from Partial Hypothesis SpaceJianhao Yan, Chenming Wu, Fandong Meng, Jie ZhouEMNLP 2022
