On the Weaknesses of Reinforcement Learning for Neural Machine Translation
Leshem Choshen, Lior Fox, Zohar Aizenbud, Omri Abend
摘要
Reinforcement learning (RL) is frequently used to increase performance in text generation tasks, including machine translation (MT), notably through the use of Minimum Risk Training (MRT) and Generative Adversarial Networks (GAN). However, little is known about what and how these methods learn in the context of MT. We prove that one of the most common RL methods for MT does not optimize the expected reward, as well as show that other methods take an infeasibly long time to converge. In fact, our results suggest that RL practices in MT are likely to improve performance only where the pre-trained parameters are already close to yielding the correct translation. Our findings further suggest that observed gains may be due to effects unrelated to the training signal, but rather from changes in the shape of the distribution curve.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper37
- Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-constraintWei Xiong, Hanze Dong, Chenlu Ye, Ziqi Wang 等ICML 2024 · 被引用 346 次
- Contrastive Learning with Adversarial Perturbations for Conditional Text GenerationSeanie Lee, Dong Bok Lee, Sung Ju HwangICLR 2021 · 被引用 117 次
- Beyond Binary Rewards: Training LMs to Reason About Their UncertaintyMehul Damani, Isha Puri, Stewart Slocum, Idan Shenfeld 等ICLR 2026 · 被引用 116 次
- On Reinforcement Learning and Distribution Matching for Fine-Tuning Language Models with no Catastrophic ForgettingTomasz Korbak, Hady Elsahar, Germán Kruszewski, Marc DymetmanNeurIPS 2022 · 被引用 98 次
- Text Generation by Learning from DemonstrationsRichard Yuanzhe Pang, He HeICLR 2021 · 被引用 88 次
它引用的顶会 Paper1
相关 Paper
- MLE-Guided Parameter Search for Task Loss Minimization in Neural Sequence ModelingSean Welleck, Kyunghyun ChoAAAI 2021 · 被引用 8 次
- ColdGANs: Taming Language GANs with Cautious Sampling StrategiesThomas Scialom, Paul-Alexis Dray, Sylvain Lamprier, Benjamin Piwowarski 等NeurIPS 2020 · 被引用 19 次
- TextGAIL: Generative Adversarial Imitation Learning for Text GenerationQingyang Wu, Lei Li, Zhou YuAAAI 2021 · 被引用 54 次
- Rephrasing the Reference for Non-autoregressive Machine TranslationChenze Shao, Jinchao Zhang, Jie Zhou, Yang FengAAAI 2023 · 被引用 6 次
- Generative Cooperative Networks for Natural Language GenerationSylvain Lamprier, Thomas Scialom, Antoine Chaffin, Vincent Claveau 等ICML 2022 · 被引用 13 次
