Enhancing Reinforcement Learning with Dense Rewards from Language Model Critic
Meng Cao, Lei Shu, Lei Yu, Yun Zhu, Nevan Wichers, Yinxiao Liu, Lei Meng
摘要
Reinforcement learning (RL) can align language models with non-differentiable reward signals, such as human preferences. However, a major challenge arises from the sparsity of these reward signals -typically, there is only a single reward for an entire output. This sparsity of rewards can lead to inefficient and unstable learning. To address this challenge, our paper introduces an novel framework that utilizes the critique capability of Large Language Models (LLMs) to produce intermediate-step rewards during RL training. Our approach pairs a policy model with a critic language model that provides feedback on each part of the policy's output. This feedback is then translated into token or span-level rewards that can be used to guide the RL training process. We investigate this approach under two different settings: one where the policy model is smaller and is paired with a more powerful critic model, and another where a single language model fulfills both roles. We assess our approach on three text generation tasks: sentiment control, language model detoxification, and summarization. Experimental results show that incorporating artificial intrinsic rewards significantly improve both sample efficiency and the overall performance of the policy model, supported by both automatic and human evaluation. The code is available under Google Research github * .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- SRPO: Self-Referential Policy Optimization for Vision-Language-Action ModelsSenyu Fei, Siyin Wang, Li Ji, Ao Li 等CVPR 2026 · 被引用 28 次
- ReDit: Reward Dithering for Improved LLM Policy OptimizationChenxing Wei, Jiarui Yu, Ying He, Hande Dong 等NeurIPS 2025 · 被引用 14 次
- Dynamic and Generalizable Process Reward ModelingZhangyue Yin, Qiushi Sun, Zhiyuan Zeng, Qinyuan Cheng 等ACL 2025 · 被引用 13 次
- AlignDistil: Token-Level Language Model Alignment as Adaptive Policy DistillationSongming Zhang, Xue Zhang, Tong Zhang, Bojie Hu 等ACL 2025
- Towards Cost-Effective Reward Guided Text GenerationAhmad Rashid, Ruotian Wu, Rongqi Fan, Hongliang Li 等ICML 2025
它引用的顶会 Paper20
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan 等NeurIPS 2023 · 被引用 4,972 次
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley 等ICML 2023 · 被引用 1,822 次
- Plug and Play Language Models: A Simple Approach to Controlled Text GenerationSumanth Dathathri, Andrea Madotto, Janice Lan, Jane Hung 等ICLR 2020 · 被引用 1,166 次
- Fine-Grained Human Feedback Gives Better Rewards for Language Model TrainingZeqiu Wu, Yushi Hu, Weijia Shi, Nouha Dziri 等NeurIPS 2023 · 被引用 516 次
相关 Paper
- Curiosity-Driven Reinforcement Learning from Human FeedbackHaoran Sun, Yekun Chai, Shuohuan Wang, Yu Sun 等ACL 2025 · 被引用 17 次
- Text2Grad: Reinforcement Learning from Natural Language FeedbackHanyang Wang, Lu Wang, Chaoyun Zhang, Tianjun Mao 等ICLR 2026 · 被引用 18 次
- Critique-RL: Training Language Models For Critiquing Through Two-Stage Reinforcement LearningZhiheng Xi, Jixuan Huang, Xin Guo, Boyang Hong 等ICLR 2026 · 被引用 4 次
- Teaching Language Models to Critique via Reinforcement LearningZhihui Xie, Jie Chen, Liyu Chen, Weichao Mao 等ICML 2025
- Advancing LLM Reasoning with Natural Language and Numerical FeedbackXiaoying Zhang, Yipeng Zhang, Hao Sun, Kaituo Feng 等ICML 2026 · 被引用 79 次
