Nudging the Boundaries of LLM Reasoning
Justin Chih-Yao Chen, Xiangyu Peng, Prafulla Kumar Choubey, Kung-Hsiang Huang, Jiaxin Zhang, Mohit Bansal, Chien-Sheng Wu
摘要
Current online reinforcement learning (RL) algorithms like GRPO share a key limitation in LLM reasoning: they cannot learn from problems that are "unsolvable" to the model. In other words, they can only improve performance on problems where the model is capable of exploring the correct answer. If a problem is too difficult -- such that even hundreds of attempts never produce a correct solution -- the model cannot learn from it. Consequently, the model's "upper limit" remains unchanged after RL training, even though the likelihood of solving easier, solvable problems may increase. These hard, unsolvable samples -- though potentially rich in learning signal -- cannot contribute to training, as no rollouts yield rewards and thus no gradients are produced. To unlock learning from these hard samples, we propose NuRL, a "nudging" method that aims to push the upper bound of LLM reasoning using self-generated hints, i.e., abstract cues that help reduce the problem difficulty for the model. Given a question and its gold answer, the model generates a Chain-of-Thought (CoT) and then produces a hint containing the core knowledge needed to solve the problem. During online RL training, we generate G rollouts from the base policy and use the pass rate to decide whether the hint should be injected. For hard samples with a 0% pass rate, we inject the offline-generated hint and regenerate a new batch of trajectories. This yields two benefits: (1) the hint boosts pass rates (from 0% to non-zero), thereby introducing training signals for previously unsolvable samples, and (2) the hints are self-generated (conditioned on the gold answer), avoiding distributional shift and do not rely on external models. Compared to standard GRPO, NuRL achieves consistent improvements across six diverse benchmarks and three models, while remaining complementary to test-time scaling. Notably, NuRL can raise the model's upper limit, whereas GRPO leaves pass@1024 unchanged from the base model. Furthermore, we present a systematic study of what makes an effective hint and when hints are most useful. Interestingly, the best hints are abstract and high-level -- as revealing gold answers actually hurt performance -- and are most beneficial when applied necessarily and after GRPO has converged.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Privileged Information Distillation for Language ModelsEmiliano Penaloza, Dheeraj Vattikonda, Nicolas Gontier, Alexandre Lacoste 等ICML 2026 · 被引用 61 次
- Maximum Likelihood Reinforcement LearningFahim Tajwar, Guanning Zeng, Yueer Zhou, Yuda Song 等ICML 2026 · 被引用 18 次
- Reuse your FLOPs: Scaling RL on Hard Problems by Conditioning on Very Off-Policy PrefixesAmrith Setlur, Zijian Wang, Andrew Cohen, Paria Rashidinejad 等ICML 2026 · 被引用 13 次
- Better, Faster: Harnessing Self-Improvement in Large Reasoning ModelsQihuang Zhong, Liang Ding, Juhua Liu, Bo Du 等ICML 2026 · 被引用 3 次
- Do Not Step Into the Same River Twice: Learning to Reason from Trial and ErrorChenming Tang, Hsiu-Yuan Huang, Weijie Liu, Clive Bai 等ACL 2026 · 被引用 2 次
它引用的顶会 Paper27
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo 等NeurIPS 2022 · 被引用 8,168 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards 等ICLR 2024 · 被引用 3,045 次
相关 Paper
- RESTRAIN: From Spurious Votes to Signals - Self-Training RL with Self-PenalizationZhaoning Yu, Zhaolun Su, Leitian Tao, Haozhu Wang 等ICLR 2026 · 被引用 7 次
- Scaf-GRPO: Scaffolded Group Relative Policy Optimization for Enhancing LLM ReasoningXichen Zhang, Sitong Wu, Yinghao Zhu, Haoru Tan 等ICLR 2026 · 被引用 52 次
- Boosting MLLM Reasoning with Text-Debiased Hint-GRPOQihan Huang, Weilong Dai, Jinlong Liu, Wanggui He 等ICCV 2025
- Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective RolloutsHaizhong Zheng, Yang Zhou, Brian R. Bartoldson, Bhavya Kailkhura 等NeurIPS 2025 · 被引用 125 次
- XRPO: Pushing the Limits of GRPO with Targeted Exploration and ExploitationUdbhav Bamba, Minghao Fang, Yifan Yu, Haizhong Zheng 等ICML 2026 · 被引用 17 次
