On the Weaknesses of Reinforcement Learning for Neural Machine Translation
Leshem Choshen, Lior Fox, Zohar Aizenbud, Omri Abend
Abstract
Reinforcement learning (RL) is frequently used to increase performance in text generation tasks, including machine translation (MT), notably through the use of Minimum Risk Training (MRT) and Generative Adversarial Networks (GAN). However, little is known about what and how these methods learn in the context of MT. We prove that one of the most common RL methods for MT does not optimize the expected reward, as well as show that other methods take an infeasibly long time to converge. In fact, our results suggest that RL practices in MT are likely to improve performance only where the pre-trained parameters are already close to yielding the correct translation. Our findings further suggest that observed gains may be due to effects unrelated to the training signal, but rather from changes in the shape of the distribution curve.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers37
- Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-constraintWei Xiong, Hanze Dong, Chenlu Ye, Ziqi Wang et al.ICML 2024 · 346 citations
- Contrastive Learning with Adversarial Perturbations for Conditional Text GenerationSeanie Lee, Dong Bok Lee, Sung Ju HwangICLR 2021 · 117 citations
- Beyond Binary Rewards: Training LMs to Reason About Their UncertaintyMehul Damani, Isha Puri, Stewart Slocum, Idan Shenfeld et al.ICLR 2026 · 116 citations
- On Reinforcement Learning and Distribution Matching for Fine-Tuning Language Models with no Catastrophic ForgettingTomasz Korbak, Hady Elsahar, Germán Kruszewski, Marc DymetmanNeurIPS 2022 · 98 citations
- Text Generation by Learning from DemonstrationsRichard Yuanzhe Pang, He HeICLR 2021 · 88 citations
Builds on1
Related papers
- MLE-Guided Parameter Search for Task Loss Minimization in Neural Sequence ModelingSean Welleck, Kyunghyun ChoAAAI 2021 · 8 citations
- ColdGANs: Taming Language GANs with Cautious Sampling StrategiesThomas Scialom, Paul-Alexis Dray, Sylvain Lamprier, Benjamin Piwowarski et al.NeurIPS 2020 · 19 citations
- TextGAIL: Generative Adversarial Imitation Learning for Text GenerationQingyang Wu, Lei Li, Zhou YuAAAI 2021 · 54 citations
- Rephrasing the Reference for Non-autoregressive Machine TranslationChenze Shao, Jinchao Zhang, Jie Zhou, Yang FengAAAI 2023 · 6 citations
- Generative Cooperative Networks for Natural Language GenerationSylvain Lamprier, Thomas Scialom, Antoine Chaffin, Vincent Claveau et al.ICML 2022 · 13 citations
