The Choice of Divergence: A Neglected Key to Mitigating Diversity Collapse in Reinforcement Learning with Verifiable Reward
Long Li, Zhijian Zhou, Jiaran Hao, Jason Klein Liu, Yanting Miao, Wei Pang, Xiaoyu Tan, Wei Chu, Zhe Wang, Shirui Pan, Chao Qu, Yuan Qi
摘要
A central paradox in fine-tuning Large Language Models (LLMs) with Reinforcement Learning with Verifiable Reward (RLVR) is the frequent degradation of multi-attempt performance (Pass@k) despite improvements in single-attempt accuracy (Pass@1). This is often accompanied by catastrophic forgetting, where models lose previously acquired skills. Despite numerous proposed methods, the community's focus on the standard reverse-KL divergence has led to a surprising oversight: the potential of alternative f-divergences as a proactive solution has been largely unexamined. We argue that standard RLVR objectives-both those using the mode-seeking reverse-KL divergence and those forgoing a divergence term entirely-lack a crucial mechanism for knowledge retention. The reverse-KL actively accelerates this decay by narrowing the policy, while its absence provides no safeguard against the model drifting from its diverse knowledge base. We propose a fundamental shift in perspective: using the divergence term itself as the solution. Our framework, Diversity-Preserving Hybrid RL (DPH-RL), leverages mass-covering f-divergences (like forward-KL and JS-divergence) to function as a 'rehearsal mechanism'. By continuously referencing the initial policy, this approach forces the model to maintain broad solution coverage. Math and SQL generation experiments show that DPH-RL surpasses the GRPO baseline by improving both in-domain Pass@1 and Pass@k scores, and effectively prevents catastrophic forgetting on out-of-domain tasks. Additionally, DPH-RL is more training-efficient because it computes f-divergence using generator functions, requiring only sampling from the initial policy and no online reference model. Our work highlights a crucial, overlooked axis for improving RLVR, demonstrating that the proper selection of a divergence measure is a powerful tool for building more general and diverse reasoning models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Information Gain-based Policy Optimization: A Simple and Effective Approach for Multi-Turn Search AgentsGuoqing Wang, Sunhao Dai, Guangze Ye, Zeyu Gan 等ICLR 2026 · 被引用 37 次
- Whatever Remains Must Be True: Filtering Drives Reasoning in LLMs, Shaping DiversityGermán Kruszewski, Pierre Erbacher, Jos Rozen, Marc DymetmanICLR 2026 · 被引用 3 次
- SAGE: Shaping Anchors for Guided Exploration in RLVR of LLMsChanuk Lee, Minki Kang, Sung Ju HwangICML 2026 · 被引用 1 次
- : Hierarchical Multi-Agent Decision-Making with LLM-Based Agents in Interactive EnvironmentsSangeun Park, Minhae KwonICML 2026
- DisPPO: Quantile-Based Distributional Reinforcement Learning for Large Language ModelsZhijian Zhou, Long Li, Xuan Zhang, Zongkai Liu 等ICML 2026
它引用的顶会 Paper12
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Solving Quantitative Reasoning Problems with Language ModelsAitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer 等NeurIPS 2022 · 被引用 2,039 次
- Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?Zhiqi Chen, Rui Lu, Andrew Zhao, Zhaokai Wang 等NeurIPS 2025 · 被引用 1,109 次
- The Unreasonable Effectiveness of Entropy Minimization in LLM ReasoningShivam Agarwal, Zimin Zhang, Lifan Yuan, Jiawei Han 等NeurIPS 2025 · 被引用 185 次
- AceReason-Nemotron 1.1: Advancing Math and Code Reasoning through SFT and RL SynergyZihan Liu, Zhuolin Yang, Yang Chen, Chankyu Lee 等ICLR 2026 · 被引用 73 次
相关 Paper
- KL-Regularized Reinforcement Learning for Generative Modelling is Designed to Mode CollapseAnthony GX-Chen, Jatin Prakash, Jeff Guo, Rob Fergus 等ICLR 2026 · 被引用 18 次
- -Divergence Regularized RLHF: Two Tales of Sampling and Unified AnalysesDi Wu, Chengshuai Shi, Jing Yang, Cong ShenICML 2026
- Risk-Sensitive Reinforcement Learning for Alleviating Exploration Dilemmas in Large Language ModelsYuhua Jiang, Jiawei Huang, Yufeng Yuan, Xin Mao 等ICLR 2026 · 被引用 8 次
- FlowRL: Matching Reward Distributions for LLM ReasoningXuekai Zhu, Daixuan Cheng, Dinghuai Zhang, Hengli Li 等ICLR 2026 · 被引用 41 次
- RL's Razor: Why Online Reinforcement Learning Forgets LessIdan Shenfeld, Jyothish Pari, Pulkit AgrawalICLR 2026 · 被引用 176 次
