Diversity-Aware Policy Optimization for Large Language Model Reasoning
Jian Yao, Ran Cheng, Xingyu Wu, Jibin Wu, Kay Chen Tan
Abstract
The reasoning capabilities of large language models (LLMs) have advanced rapidly, particularly following the release of DeepSeek-R1, which has inspired a surge of research into data quality and reinforcement learning (RL) algorithms. Despite the pivotal role diversity plays in RL, its influence on LLM reasoning remains largely underexplored. To bridge this gap, this work presents a systematic investigation into the impact of diversity in RL-based training for LLM reasoning, and proposes a novel diversity-aware policy optimization method. Across evaluations on 12 LLMs, we observe a strong positive correlation between the solution diversity and Potential@k (a novel metric quantifying an LLM's reasoning potential) in high-performing models. This finding motivates our method to explicitly promote diversity during RL training. Specifically, we design a token-level diversity and reformulate it into a practical objective, then we selectively apply it to positive samples. Integrated into the R1-zero training framework, our method achieves a 3.5% average improvement across four mathematical reasoning benchmarks, while generating more diverse and robust solutions. The code is available at https://github.com/nigelyaoj/R1_zero_Div.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5cedefc3-7f01-41a0-b69c-7cc0e71d164cCited by top-tier papers15
- Diversity-Incentivized Exploration for Versatile ReasoningZican Hu, Shilin Zhang, Yafu Li, Jianhao Yan et al.ICLR 2026 · 32 citations
- Exploration vs Exploitation: Rethinking RLVR through Clipping, Entropy, and Spurious RewardPeter Chen, Xiaopeng Li, Ziniu Li, Wotao Yin et al.ICLR 2026 · 28 citations
- Post-training Large Language Models for Diverse High-Quality ResponsesYilei Chen, Souradip Chakraborty, Lorenz Wolf, Ioannis Paschalidis et al.ICLR 2026 · 20 citations
- KL-Regularized Reinforcement Learning for Generative Modelling is Designed to Mode CollapseAnthony GX-Chen, Jatin Prakash, Jeff Guo, Rob Fergus et al.ICLR 2026 · 18 citations
- HM3: Hierarchical Multi-Objective Model Merging for Pretrained ModelsYu Zhou, Xingyu Wu, Jibin Wu, Liang Feng et al.NeurIPS 2025 · 14 citations
Builds on25
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards et al.ICLR 2024 · 3,045 citations
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan et al.NeurIPS 2025 · 2,828 citations
- Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?Zhiqi Chen, Rui Lu, Andrew Zhao, Zhaokai Wang et al.NeurIPS 2025 · 1,109 citations
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
- Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base ModelJingcheng Hu, Yinmin Zhang, Qi Han, Daxin Jiang et al.NeurIPS 2025 · 533 citations
Related papers
- General-Reasoner: Advancing LLM Reasoning Across All DomainsXueguang Ma, Qian Liu, Dongfu Jiang, Ge Zhang et al.NeurIPS 2025 · 153 citations
- Is PRM Necessary? Problem-Solving RL Implicitly Induces PRM Capability in LLMsZhangyin Feng, Qianglong Chen, Ning Lu, Yongqian Li et al.NeurIPS 2025 · 16 citations
- SetPO: Set-Level Policy Optimization for Diversity-Preserving LLM ReasoningChenyi Li, Yuan Zhang, Bo Wang, Guoqing Ma et al.ICML 2026 · 4 citations
- Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMsXumeng Wen, Zihan Liu, Shun Zheng, Shengyu Ye et al.ICLR 2026 · 279 citations
- Diversity-Enhanced Reasoning for Subjective QuestionsYumeng Wang, Zhiyuan Fan, Jiayu Liu, Jen-Tse Huang et al.ICLR 2026 · 13 citations
