Replay Failures as Successes: Sample-Efficient Reinforcement Learning for Instruction Following
Kongcheng Zhang, QI YAO, Shunyu Liu, Wenjian Zhang, Min Cen, Yang Zhou, Wenkai Fang, Yiru Zhao, Baisheng Lai, Mingli Song
Abstract
Reinforcement Learning (RL) has shown promise for aligning Large Language Models (LLMs) to follow instructions with various constraints. Despite the encouraging results, RL improvement inevitably relies on sampling successful, high-quality responses; however, the initial model often struggles to generate responses that satisfy all constraints due to its limited capabilities, yielding sparse or indistinguishable rewards that impede learning. In this work, we propose ** H indsight i nstruction R eplay (HiR), a novel sample-efficient RL framework for complex instruction following tasks, which employs a select -then- rewrite strategy to replay failed attempts as successes based on the constraints that have been satisfied in hindsight. We perform RL on these replayed samples as well as the original ones, theoretically framing the objective as dual-preference learning at both the instruction- and response-level to enable efficient optimization using only a binary reward signal. Extensive experiments demonstrate that the proposed HiR yields promising results across different instruction following tasks, while requiring less computational budget. Our code and dataset are available at https://github.com/sastpg/HIR.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cbb4cf77-bbb6-4432-abf9-495e061425f0Builds on16
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards et al.ICLR 2024 · 3,045 citations
- PaLM-E: An Embodied Multimodal Language ModelDanny Driess, Fei Xia, Mehdi S. M. Sajjadi, Corey Lynch et al.ICML 2023 · 2,601 citations
- Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?Zhiqi Chen, Rui Lu, Andrew Zhao, Zhaokai Wang et al.NeurIPS 2025 · 1,109 citations
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
Related papers
- Experience is the Best Teacher: Motivating Effective Exploration in Reinforcement Learning for LLMsWenjian Zhang, Kongcheng Zhang, Jiaxin Qi, Jianqiang Huang et al.ICML 2026 · 1 citation
- The Wisdom of Hindsight Makes Language Models Better Instruction FollowersTianjun Zhang, Fangchen Liu, Justin Wong, Pieter Abbeel et al.ICML 2023 · 65 citations
- Towards Better Correctness and Efficiency in Code GenerationYunlong Feng, Yang Xu, Xiao Xu, Binyuan Hui et al.AAAI 2026 · 3 citations
- DecIF: Improving Instruction-Following through DecompositionTingfeng Hui, Pengyu Zhu, Bowen Ping, Ling Tang et al.ACL 2026
- LARFT: Closing the Cognition-Action Gap for Length Instruction Following in Large Language ModelsWei Zhang, Lintong Du, yuanhe zhang, Zhenhong Zhou et al.ICML 2026
