XRPO: Pushing the Limits of GRPO with Targeted Exploration and Exploitation
Udbhav Bamba, Minghao Fang, Yifan Yu, Haizhong Zheng, Fan Lai
Abstract
Reinforcement learning algorithms such as GRPO have driven recent advances in large language model (LLM) reasoning. While scaling the number of rollouts stabilizes training, existing approaches suffer from limited exploration on challenging prompts and leave informative feedback signals underexploited, due to context-independent rollout allocation across prompts (e.g., generating 16 rollouts per prompt) and relying heavily on sparse rewards. This paper presents XRPO (eXplore–eXploit GRPO), a unified framework that recasts policy optimization through the principled lens of rollout exploration–exploitation. To enhance exploration, XRPO introduces a mathematically grounded rollout allocator that adaptively prioritizes prompts with higher potential for uncertainty reduction. It further addresses stagnation on zero-reward prompts through an in-context seeding strategy that injects curated exemplars, steering the model into more difficult reasoning trajectories. To strengthen exploitation, XRPO develops a group-relative, novelty-aware advantage sharpening mechanism that leverages sequence likelihoods to amplify low-probability yet correct responses, thereby extending the policy’s reach beyond sparse rewards. Experiments across diverse math and coding benchmarks on both reasoning and non-reasoning models demonstrate that XRPO outperforms existing advances (e.g., GRPO and GSPO) up to 4% pass@1 and 6% cons@32, while accelerating training convergence by up to 2.7x.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6f4d25dd-2717-4a76-9a10-5634af32dc3aCited by top-tier papers4
- CodeV: Code with Images for Faithful Visual Reasoning via Tool-Aware Policy OptimizationXinhai Hou, Shaoyuan Xu, Manan Biyani, Moyan Li et al.CVPR 2026 · 25 citations
- RLBoost: Harvesting Preemptible Cloud Resources for Cost-Efficient Reinforcement Learning on LLMsYongji Wu, Xueshen Liu, Haizhong Zheng, Juncheng Gu et al.NSDI 2026 · 4 citations
- QUATRO: Query-Adaptive Trust Region Policy Optimization for LLM Fine-tuningDoyeon Lee, Eunyi Lyou, Hyunsoo Cho, Soo Kyung Kim et al.ICML 2026
- Miner: Mining Intrinsic Mastery for Data-Efficient RL in Large Reasoning ModelsShuyang Jiang, Yuhao Wang, Ya Zhang, Yanfeng Wang et al.ACL 2026
Builds on16
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards et al.ICLR 2024 · 3,045 citations
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan et al.NeurIPS 2025 · 2,828 citations
- Teaching Large Language Models to Self-DebugXinyun Chen, Maxwell Lin, Nathanael Schärli, Denny ZhouICLR 2024 · 1,085 citations
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
- Large Language Models Cannot Self-Correct Reasoning YetJie Huang, Xinyun Chen, Swaroop Mishra, Huaixiu Steven Zheng et al.ICLR 2024 · 858 citations
Related papers
- Slow-Fast Policy Optimization: Reposition-Before-Update for LLM ReasoningZiyan Wang, Zheng Wang, Xingwei Qu, Qi Cheng et al.ICLR 2026 · 4 citations
- Knapsack RL: Compute-Efficient Reinforcement Learning via Heterogeneous Rollout AllocationZiniu Li, Congliang Chen, Tianyun Yang, Tian Ding et al.ICML 2026
- ExGRPO: Learning to Reason from ExperienceRunzhe Zhan, Yafu Li, Zhi Wang, Xiaoye Qu et al.ICLR 2026 · 51 citations
- Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective RolloutsHaizhong Zheng, Yang Zhou, Brian R. Bartoldson, Bhavya Kailkhura et al.NeurIPS 2025 · 125 citations
- Lookahead Tree-Based Rollouts for Enhanced Trajectory-Level Exploration in Reinforcement Learning with Verifiable RewardsShangyu Xing, Siyuan Wang, Chenyuan Yang, Xin-Yu Dai et al.ICLR 2026 · 14 citations
