Teacher Forcing Recovers Reward Functions for Text Generation
Yongchang Hao, Yuxin Liu, Lili Mou
Abstract
Reinforcement learning (RL) has been widely used in text generation to alleviate the exposure bias issue or to utilize non-parallel datasets. The reward function plays an important role in making RL training successful. However, previous reward functions are typically task-specific and sparse, restricting the use of RL. In our work, we propose a task-agnostic approach that derives a step-wise reward function directly from a model trained with teacher forcing. We additionally propose a simple modification to stabilize the RL training on non-parallel datasets with our induced reward function. Empirical results show that our method outperforms selftraining and reward regression methods on several text generation tasks, confirming the effectiveness of our reward function. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8ea12187-619d-4eb1-8467-a86a32a11142Cited by top-tier papers4
- MiniLLM: Knowledge Distillation of Large Language ModelsYuxian Gu, Li Dong, Furu Wei, Minlie HuangICLR 2024 · 95 citations
- Towards Brain Passage Retrieval: An Investigation of EEG Query RepresentationsNiall McGuire, Yashar MoshfeghiSIGIR 2025 · 5 citations
- Adversarial Moment-Matching Distillation of Large Language ModelsChen JiaNeurIPS 2024 · 4 citations
- Grammar-Forced Translation of Natural Language to Temporal Logic using LLMsWilliam H. English, Dominic Simon, Sumit Kumar Jha, Rickard EwetzICML 2025
Builds on14
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- SQIL: Imitation Learning via Reinforcement Learning with Sparse RewardsSiddharth Reddy, Anca D. Dragan, Sergey LevineICLR 2020 · 299 citations
- Text Generation by Learning from DemonstrationsRichard Yuanzhe Pang, He HeICLR 2021 · 88 citations
- DialogBERT: Discourse-Aware Response Generation via Learning to Recover and Rank UtterancesXiaodong Gu, Kang Min Yoo, Jung-Woo HaAAAI 2021 · 83 citations
- Unsupervised Paraphrasing by Simulated AnnealingXianggen Liu, Lili Mou, Fandong Meng, Hao Zhou et al.ACL 2020 · 74 citations
Related papers
- GRAM: A Generative Foundation Reward Model for Reward GeneralizationChenglong Wang, Yang Gan, Yifu Huo, Yongyu Mu et al.ICML 2025
- On the Effectiveness of Offline RL for Dialogue Response GenerationPaloma Sodhi, Felix Wu, Ethan R. Elenberg, Kilian Q. Weinberger et al.ICML 2023 · 6 citations
- Fantastic Rewards and How to Tame Them: A Case Study on Reward Learning for Task-oriented Dialogue SystemsYihao Feng, Shentao Yang, Shujian Zhang, Jianguo Zhang et al.ICLR 2023 · 6 citations
- Text2Reward: Reward Shaping with Language Models for Reinforcement LearningTianbao Xie, Siheng Zhao, Chen Henry Wu, Yitao Liu et al.ICLR 2024 · 142 citations
- Towards Cost-Effective Reward Guided Text GenerationAhmad Rashid, Ruotian Wu, Rongqi Fan, Hongliang Li et al.ICML 2025
