Rewarding the Unlikely: Lifting GRPO Beyond Distribution Sharpening
Andre Wang He, Daniel Fried, Sean Welleck
摘要
Reinforcement learning is emerging as a primary driver for improving language model reasoning capabilities. A fundamental question is whether current reinforcement learning algorithms -- such as Group Relative Policy Optimization (GRPO), the de facto standard algorithm used to improve language model reasoning -- merely sharpen the base model's distribution around problems it can already solve. We investigate this question in the context of formal theorem proving, which has access to a perfect verifier. We identify a degenerate rank bias in GRPO in which highly probable trajectories are reinforced and rare ones are neglected. This results in distribution sharpening: the model can solve some problems with fewer samples, but underperforms simply sampling more solutions from the original model. To overcome GRPO's rank bias we introduce unlikeliness reward, a simple method for explicitly up-weighting rare but correct solutions. We show that unlikeliness reward mitigates rank bias and improves pass@ across a large range of in both synthetic and real theorem proving settings. We also uncover an unexpected link between rank bias and a seemingly mundane hyperparameter -- the number of updates per batch -- that leads to a second, complementary mitigation. We combine our insights into a revised GRPO training recipe for formal theorem proving, yielding an open pipeline that achieves competitive performance to DeepSeek-Prover-V1.5-RL on the miniF2F-test benchmark. We release our implementation at https://github.com/AndreHe02/rewarding-unlikely-release
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper19
- Reasoning with Sampling: Your Base Model is Smarter Than You ThinkAayush Karan, Yilun DuICLR 2026 · 被引用 87 次
- From f(x) and g(x) to f(g(x)): LLMs Learn New Skills in RL by Composing Old OnesLifan Yuan, Weize Chen, Yuchen Zhang, Ganqu Cui 等ICLR 2026 · 被引用 46 次
- Nudging the Boundaries of LLM ReasoningJustin Chih-Yao Chen, Xiangyu Peng, Prafulla Kumar Choubey, Kung-Hsiang Huang 等ICLR 2026 · 被引用 25 次
- Post-training Large Language Models for Diverse High-Quality ResponsesYilei Chen, Souradip Chakraborty, Lorenz Wolf, Ioannis Paschalidis 等ICLR 2026 · 被引用 20 次
- KL-Regularized Reinforcement Learning for Generative Modelling is Designed to Mode CollapseAnthony GX-Chen, Jatin Prakash, Jeff Guo, Rob Fergus 等ICLR 2026 · 被引用 18 次
它引用的顶会 Paper11
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- miniF2F: a cross-system benchmark for formal Olympiad-level mathematicsKunhao Zheng, Jesse Michael Han, Stanislas PoluICLR 2022 · 被引用 342 次
- HyperTree Proof Search for Neural Theorem ProvingGuillaume Lample, Timothée Lacroix, Marie-Anne Lachaux, Aurélien Rodriguez 等NeurIPS 2022 · 被引用 271 次
- BFS-Prover: Scalable Best-First Tree Search for LLM-based Automatic Theorem ProvingRan Xin, Chenguang Xi, Jie Yang, Feng Chen 等ACL 2025 · 被引用 66 次
- HybridFlow: A Flexible and Efficient RLHF FrameworkGuangming Sheng, Chi Zhang, Zilingfeng Ye, Xibin Wu 等EuroSys 2025 · 被引用 61 次
相关 Paper
- On the Effect of Negative Gradient in Group Relative Deep Reinforcement OptimizationWenlong Deng, Yi Ren, Muchen Li, Danica J. Sutherland 等NeurIPS 2025 · 被引用 36 次
- Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMsZhihe Yang, Xufang Luo, Zilong Wang, Dongqi Han 等ICLR 2026 · 被引用 47 次
- Reinforcement Learning with Verifiable Rewards: GRPO's Loss, Dynamics, and Success AmplificationYoussef MrouehICML 2026 · 被引用 118 次
- XRPO: Pushing the Limits of GRPO with Targeted Exploration and ExploitationUdbhav Bamba, Minghao Fang, Yifan Yu, Haizhong Zheng 等ICML 2026 · 被引用 17 次
- Token Hidden Reward: Steering Exploration-Exploitation in Group Relative Deep Reinforcement LearningWenlong Deng, Yi Ren, Yushu Li, Boying Gong 等ICLR 2026 · 被引用 9 次
