Text Generation by Learning from Demonstrations
Richard Yuanzhe Pang, He He
Abstract
Current approaches to text generation largely rely on autoregressive models and maximum likelihood estimation. This paradigm leads to (i) diverse but low-quality samples due to mismatched learning objective and evaluation metric (likelihood vs. quality) and (ii) exposure bias due to mismatched history distributions (gold vs. model-generated). To alleviate these problems, we frame text generation as an offline reinforcement learning (RL) problem with expert demonstrations (i.e., the reference), where the goal is to maximize quality given model-generated histories. We propose GOLD (generation by off-policy learning from demonstrations): an easy-to-optimize algorithm that learns from the demonstrations by importance weighting. Intuitively, GOLD upweights confident tokens and downweights unconfident ones in the reference during training, avoiding optimization issues faced by prior RL approaches that rely on online data collection. According to both automatic and human evaluation, models trained by GOLD outperform those trained by MLE and policy gradient on summarization, question generation, and machine translation. Further, our models are less sensitive to decoding algorithms and alleviate exposure bias.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 68506d28-a4e1-451a-a060-c6b02f6b87b2Cited by top-tier papers35
- BRIO: Bringing Order to Abstractive SummarizationYixin Liu, Pengfei Liu, Dragomir R. Radev, Graham NeubigACL 2022 · 329 citations
- Coarse-to-Fine Vision-Language Pre-training with Fusion in the BackboneZi-Yi Dou, Aishwarya Kamath, Zhe Gan, Pengchuan Zhang et al.NeurIPS 2022 · 173 citations
- On the Generalization of SFT: A Reinforcement Learning Perspective with Reward RectificationYongliang Wu, Yizhou Zhou, Ziheng Zhou, Yingzhe Peng et al.ICLR 2026 · 130 citations
- MiniLLM: Knowledge Distillation of Large Language ModelsYuxian Gu, Li Dong, Furu Wei, Minlie HuangICLR 2024 · 95 citations
- Is Reinforcement Learning (Not) for Natural Language Processing: Benchmarks, Baselines, and Building Blocks for Natural Language Policy OptimizationRajkumar Ramamurthy, Prithviraj Ammanabrolu, Kianté Brantley, Jack Hessel et al.ICLR 2023 · 54 citations
Builds on8
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive SummarizationJingqing Zhang, Yao Zhao, Mohammad Saleh, Peter J. LiuICML 2020 · 2,453 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- Neural Text Generation With Unlikelihood TrainingSean Welleck, Ilia Kulikov, Stephen Roller, Emily Dinan et al.ICLR 2020 · 683 citations
- Language GANs Falling ShortMassimo Caccia, Lucas Caccia, William Fedus, Hugo Larochelle et al.ICLR 2020 · 236 citations
Related papers
- ColdGANs: Taming Language GANs with Cautious Sampling StrategiesThomas Scialom, Paul-Alexis Dray, Sylvain Lamprier, Benjamin Piwowarski et al.NeurIPS 2020 · 19 citations
- Offline RL by Reward-Weighted Fine-Tuning for Conversation OptimizationSubhojyoti Mukherjee, Viet Dac Lai, Raghavendra Addanki, Ryan Rossi et al.NeurIPS 2025 · 12 citations
- Adaptive Prior-Dependent Correction Enhanced Reinforcement Learning for Natural Language GenerationWei Cheng, Ziyan Luo, Qiyue YinAAAI 2021 · 1 citation
- KRLS: Improving End-to-End Response Generation in Task Oriented Dialog with Reinforced Keywords LearningXiao Yu, Qingyang Wu, Kun Qian, Zhou YuEMNLP 2023 · 1 citation
- Improving Text Generation with Student-Forcing Optimal TransportJianqiao Li, Chunyuan Li, Guoyin Wang, Hao Fu et al.EMNLP 2020 · 11 citations
