Teacher Forcing Recovers Reward Functions for Text Generation
Yongchang Hao, Yuxin Liu, Lili Mou
摘要
Reinforcement learning (RL) has been widely used in text generation to alleviate the exposure bias issue or to utilize non-parallel datasets. The reward function plays an important role in making RL training successful. However, previous reward functions are typically task-specific and sparse, restricting the use of RL. In our work, we propose a task-agnostic approach that derives a step-wise reward function directly from a model trained with teacher forcing. We additionally propose a simple modification to stabilize the RL training on non-parallel datasets with our induced reward function. Empirical results show that our method outperforms selftraining and reward regression methods on several text generation tasks, confirming the effectiveness of our reward function. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- MiniLLM: Knowledge Distillation of Large Language ModelsYuxian Gu, Li Dong, Furu Wei, Minlie HuangICLR 2024 · 被引用 95 次
- Towards Brain Passage Retrieval: An Investigation of EEG Query RepresentationsNiall McGuire, Yashar MoshfeghiSIGIR 2025 · 被引用 5 次
- Adversarial Moment-Matching Distillation of Large Language ModelsChen JiaNeurIPS 2024 · 被引用 4 次
- Grammar-Forced Translation of Natural Language to Temporal Logic using LLMsWilliam H. English, Dominic Simon, Sumit Kumar Jha, Rickard EwetzICML 2025
它引用的顶会 Paper14
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- SQIL: Imitation Learning via Reinforcement Learning with Sparse RewardsSiddharth Reddy, Anca D. Dragan, Sergey LevineICLR 2020 · 被引用 299 次
- Text Generation by Learning from DemonstrationsRichard Yuanzhe Pang, He HeICLR 2021 · 被引用 88 次
- DialogBERT: Discourse-Aware Response Generation via Learning to Recover and Rank UtterancesXiaodong Gu, Kang Min Yoo, Jung-Woo HaAAAI 2021 · 被引用 83 次
- Unsupervised Paraphrasing by Simulated AnnealingXianggen Liu, Lili Mou, Fandong Meng, Hao Zhou 等ACL 2020 · 被引用 74 次
相关 Paper
- GRAM: A Generative Foundation Reward Model for Reward GeneralizationChenglong Wang, Yang Gan, Yifu Huo, Yongyu Mu 等ICML 2025
- On the Effectiveness of Offline RL for Dialogue Response GenerationPaloma Sodhi, Felix Wu, Ethan R. Elenberg, Kilian Q. Weinberger 等ICML 2023 · 被引用 6 次
- Fantastic Rewards and How to Tame Them: A Case Study on Reward Learning for Task-oriented Dialogue SystemsYihao Feng, Shentao Yang, Shujian Zhang, Jianguo Zhang 等ICLR 2023 · 被引用 6 次
- Text2Reward: Reward Shaping with Language Models for Reinforcement LearningTianbao Xie, Siheng Zhao, Chen Henry Wu, Yitao Liu 等ICLR 2024 · 被引用 142 次
- Towards Cost-Effective Reward Guided Text GenerationAhmad Rashid, Ruotian Wu, Rongqi Fan, Hongliang Li 等ICML 2025
