Replay Failures as Successes: Sample-Efficient Reinforcement Learning for Instruction Following
Kongcheng Zhang, QI YAO, Shunyu Liu, Wenjian Zhang, Min Cen, Yang Zhou, Wenkai Fang, Yiru Zhao, Baisheng Lai, Mingli Song
摘要
Reinforcement Learning (RL) has shown promise for aligning Large Language Models (LLMs) to follow instructions with various constraints. Despite the encouraging results, RL improvement inevitably relies on sampling successful, high-quality responses; however, the initial model often struggles to generate responses that satisfy all constraints due to its limited capabilities, yielding sparse or indistinguishable rewards that impede learning. In this work, we propose ** H indsight i nstruction R eplay (HiR), a novel sample-efficient RL framework for complex instruction following tasks, which employs a select -then- rewrite strategy to replay failed attempts as successes based on the constraints that have been satisfied in hindsight. We perform RL on these replayed samples as well as the original ones, theoretically framing the objective as dual-preference learning at both the instruction- and response-level to enable efficient optimization using only a binary reward signal. Extensive experiments demonstrate that the proposed HiR yields promising results across different instruction following tasks, while requiring less computational budget. Our code and dataset are available at https://github.com/sastpg/HIR.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper16
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards 等ICLR 2024 · 被引用 3,045 次
- PaLM-E: An Embodied Multimodal Language ModelDanny Driess, Fei Xia, Mehdi S. M. Sajjadi, Corey Lynch 等ICML 2023 · 被引用 2,601 次
- Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?Zhiqi Chen, Rui Lu, Andrew Zhao, Zhaokai Wang 等NeurIPS 2025 · 被引用 1,109 次
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
相关 Paper
- Experience is the Best Teacher: Motivating Effective Exploration in Reinforcement Learning for LLMsWenjian Zhang, Kongcheng Zhang, Jiaxin Qi, Jianqiang Huang 等ICML 2026 · 被引用 1 次
- The Wisdom of Hindsight Makes Language Models Better Instruction FollowersTianjun Zhang, Fangchen Liu, Justin Wong, Pieter Abbeel 等ICML 2023 · 被引用 65 次
- Towards Better Correctness and Efficiency in Code GenerationYunlong Feng, Yang Xu, Xiao Xu, Binyuan Hui 等AAAI 2026 · 被引用 3 次
- DecIF: Improving Instruction-Following through DecompositionTingfeng Hui, Pengyu Zhu, Bowen Ping, Ling Tang 等ACL 2026
- LARFT: Closing the Cognition-Action Gap for Length Instruction Following in Large Language ModelsWei Zhang, Lintong Du, yuanhe zhang, Zhenhong Zhou 等ICML 2026
