Machine Translation Decoding beyond Beam Search
Rémi Leblond, Jean-Baptiste Alayrac, Laurent Sifre, Miruna Pislar, Jean-Baptiste Lespiau, Ioannis Antonoglou, Karen Simonyan, Oriol Vinyals
Abstract
Beam search is the go-to method for decoding auto-regressive machine translation models. While it yields consistent improvements in terms of BLEU, it is only concerned with finding outputs with high model likelihood, and is thus agnostic to whatever end metric or score practitioners care about. Our aim is to establish whether beam search can be replaced by a more powerful metric-driven search technique. To this end, we explore numerous decoding algorithms, including some which rely on a value function parameterised by a neural network, and report results on a variety of metrics. Notably, we introduce a Monte-Carlo Tree Search (MCTS) based method and showcase its competitiveness. We provide a blueprint for how to use MCTS fruitfully in language applications, which opens promising future directions. We find that which algorithm is best heavily depends on the characteristics of the goal metric; we believe that our extensive experiments and analysis will inform further research in this area.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a91c53e1-f419-4e6c-8036-d6a1d6eab710Cited by top-tier papers23
- AlphaZero-Like Tree-Search can Guide Large Language Model Decoding and TrainingZiyu Wan, Xidong Feng, Muning Wen, Stephen Marcus McAleer et al.ICML 2024 · 325 citations
- Derivative-Free Guidance in Continuous and Discrete Diffusion Models with Soft Value-based DecodingXiner Li, Yulai Zhao, Chenyu Wang, Gabriele Scalia et al.NeurIPS 2025 · 147 citations
- Transformer-based Planning for Symbolic RegressionParshin Shojaee, Kazem Meidani, Amir Barati Farimani, Chandan K. ReddyNeurIPS 2023 · 116 citations
- Grounded Decoding: Guiding Text Generation with Grounded Models for Embodied AgentsWenlong Huang, Fei Xia, Dhruv Shah, Danny Driess et al.NeurIPS 2023 · 102 citations
- Tuning Computer Vision Models With Task RewardsAndré Susano Pinto, Alexander Kolesnikov, Yuge Shi, Lucas Beyer et al.ICML 2023 · 56 citations
Builds on6
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- On the Weaknesses of Reinforcement Learning for Neural Machine TranslationLeshem Choshen, Lior Fox, Zohar Aizenbud, Omri AbendICLR 2020 · 124 citations
- Consistency of a Recurrent Language Model With Respect to Incomplete DecodingSean Welleck, Ilia Kulikov, Jaedeok Kim, Richard Yuanzhe Pang et al.EMNLP 2020 · 37 citations
- If beam search is the answer, what was the question?Clara Meister, Ryan Cotterell, Tim VieiraEMNLP 2020 · 26 citations
Related papers
- Enabling Arbitrary Translation Objectives with Adaptive Tree SearchWang Ling, Wojciech Stokowiec, Domenic Donato, Chris Dyer et al.ICLR 2022
- Aligning Tree-Search Policies with Fixed Token Budgets in Test-Time Scaling of LLMsSora Miyamoto, Daisuke Oba, Naoaki OkazakiICML 2026 · 3 citations
- Energy-Based Reranking: Improving Neural Machine Translation Using Energy-Based ModelsSumanta Bhattacharyya, Amirmohammad Rooshenas, Subhajit Naskar, Simeng Sun et al.ACL 2021
- Non-Monotonic Latent Alignments for CTC-Based Non-Autoregressive Machine TranslationChenze Shao, Yang FengNeurIPS 2022 · 26 citations
- Latent-Variable Non-Autoregressive Neural Machine Translation with Deterministic Inference Using a Delta PosteriorRaphael Shu, Jason Lee, Hideki Nakayama, Kyunghyun ChoAAAI 2020 · 125 citations
