Rewarding the Unlikely: Lifting GRPO Beyond Distribution Sharpening
Andre Wang He, Daniel Fried, Sean Welleck
Abstract
Reinforcement learning is emerging as a primary driver for improving language model reasoning capabilities. A fundamental question is whether current reinforcement learning algorithms -- such as Group Relative Policy Optimization (GRPO), the de facto standard algorithm used to improve language model reasoning -- merely sharpen the base model's distribution around problems it can already solve. We investigate this question in the context of formal theorem proving, which has access to a perfect verifier. We identify a degenerate rank bias in GRPO in which highly probable trajectories are reinforced and rare ones are neglected. This results in distribution sharpening: the model can solve some problems with fewer samples, but underperforms simply sampling more solutions from the original model. To overcome GRPO's rank bias we introduce unlikeliness reward, a simple method for explicitly up-weighting rare but correct solutions. We show that unlikeliness reward mitigates rank bias and improves pass@ across a large range of in both synthetic and real theorem proving settings. We also uncover an unexpected link between rank bias and a seemingly mundane hyperparameter -- the number of updates per batch -- that leads to a second, complementary mitigation. We combine our insights into a revised GRPO training recipe for formal theorem proving, yielding an open pipeline that achieves competitive performance to DeepSeek-Prover-V1.5-RL on the miniF2F-test benchmark. We release our implementation at https://github.com/AndreHe02/rewarding-unlikely-release
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers19
- Reasoning with Sampling: Your Base Model is Smarter Than You ThinkAayush Karan, Yilun DuICLR 2026 · 87 citations
- From f(x) and g(x) to f(g(x)): LLMs Learn New Skills in RL by Composing Old OnesLifan Yuan, Weize Chen, Yuchen Zhang, Ganqu Cui et al.ICLR 2026 · 46 citations
- Nudging the Boundaries of LLM ReasoningJustin Chih-Yao Chen, Xiangyu Peng, Prafulla Kumar Choubey, Kung-Hsiang Huang et al.ICLR 2026 · 25 citations
- Post-training Large Language Models for Diverse High-Quality ResponsesYilei Chen, Souradip Chakraborty, Lorenz Wolf, Ioannis Paschalidis et al.ICLR 2026 · 20 citations
- KL-Regularized Reinforcement Learning for Generative Modelling is Designed to Mode CollapseAnthony GX-Chen, Jatin Prakash, Jeff Guo, Rob Fergus et al.ICLR 2026 · 18 citations
Builds on11
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- miniF2F: a cross-system benchmark for formal Olympiad-level mathematicsKunhao Zheng, Jesse Michael Han, Stanislas PoluICLR 2022 · 342 citations
- HyperTree Proof Search for Neural Theorem ProvingGuillaume Lample, Timothée Lacroix, Marie-Anne Lachaux, Aurélien Rodriguez et al.NeurIPS 2022 · 271 citations
- BFS-Prover: Scalable Best-First Tree Search for LLM-based Automatic Theorem ProvingRan Xin, Chenguang Xi, Jie Yang, Feng Chen et al.ACL 2025 · 66 citations
- HybridFlow: A Flexible and Efficient RLHF FrameworkGuangming Sheng, Chi Zhang, Zilingfeng Ye, Xibin Wu et al.EuroSys 2025 · 61 citations
Related papers
- On the Effect of Negative Gradient in Group Relative Deep Reinforcement OptimizationWenlong Deng, Yi Ren, Muchen Li, Danica J. Sutherland et al.NeurIPS 2025 · 36 citations
- Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMsZhihe Yang, Xufang Luo, Zilong Wang, Dongqi Han et al.ICLR 2026 · 47 citations
- Reinforcement Learning with Verifiable Rewards: GRPO's Loss, Dynamics, and Success AmplificationYoussef MrouehICML 2026 · 118 citations
- XRPO: Pushing the Limits of GRPO with Targeted Exploration and ExploitationUdbhav Bamba, Minghao Fang, Yifan Yu, Haizhong Zheng et al.ICML 2026 · 17 citations
- Token Hidden Reward: Steering Exploration-Exploitation in Group Relative Deep Reinforcement LearningWenlong Deng, Yi Ren, Yushu Li, Boying Gong et al.ICLR 2026 · 9 citations
