Beyond Pass@ 1: Self-Play with Variational Problem Synthesis Sustains RLVR
Xiao Liang, Zhong-Zhi Li, Yeyun Gong, Yelong Shen, Yingnian Wu, Zhijiang Guo, Weizhu Chen
摘要
˚Equal contribution, work done during internships at Microsoft. : Corresponding authors Reinforcement Learning with Verifiable Rewards (RLVR) has recently emerged as a key paradigm for post-training Large Language Models (LLMs), particularly for complex reasoning tasks. However, standard RLVR training has been shown to improve Pass@1 performance at the expense of policy entropy, leading to reduced generation diversity and limiting the Pass@k performance, which typically represents the upper bound of LLM reasoning capability. In this paper, we systematically analyze the policy's generation diversity from the perspective of training data and find that augmenting and updating training problems helps mitigate entropy collapse during training. Based on these observations, we propose an online Self-play with Variational problem Synthesis (SvS) strategy for RLVR training, which uses the policy's correct solutions to synthesize variational problems while ensuring their reference answers remain identical to the originals. This self-improving strategy effectively preserves policy entropy during training and substantially improves Pass@k compared with standard RLVR, sustaining long-term improvements and achieving absolute gains of 18.3% and 22.8% in Pass@32 performance on the competition-level AIME 24 and AIME 25 benchmarks, as well as on code generation tasks. Experiments on 12 reasoning benchmarks across varying model sizes from 3B to 32B consistently demonstrate the generalizability and robustness of SvS.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- The Choice of Divergence: A Neglected Key to Mitigating Diversity Collapse in Reinforcement Learning with Verifiable RewardLong Li, Zhijian Zhou, Jiaran Hao, Jason Klein Liu 等ICLR 2026 · 被引用 46 次
- Depth-Breadth Synergy in RLVR: Unlocking LLM Reasoning Gains with Adaptive ExplorationZhicheng Yang, Zhijiang Guo, Yinya Huang, Yongxin Wang 等ICML 2026 · 被引用 38 次
- Search Self-Play: Pushing the Frontier of Agent Capability without SupervisionHongliang Lu, Yuhang Wen, Pengyu Cheng, Ruijin Ding 等ICLR 2026 · 被引用 35 次
- Harder Is Better: Boosting Mathematical Reasoning via Difficulty-Aware GRPO and Multi-Aspect Question ReformulationYanqi Dai, Yuxiang Ji, Xiao Zhang, Yong Wang 等ICLR 2026 · 被引用 26 次
- LoopTool: Closing the Data-Training Loop for Robust LLM Tool CallsKangning Zhang, Weiwen Liu, Wenxiang Jiao, Kounianhua Du 等ACL 2026 · 被引用 18 次
它引用的顶会 Paper24
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards 等ICLR 2024 · 被引用 3,045 次
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan 等NeurIPS 2025 · 被引用 2,828 次
- Solving Quantitative Reasoning Problems with Language ModelsAitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer 等NeurIPS 2022 · 被引用 2,039 次
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
- Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base ModelJingcheng Hu, Yinmin Zhang, Qi Han, Daxin Jiang 等NeurIPS 2025 · 被引用 533 次
相关 Paper
- Rethinking Entropy Interventions in RLVR: An Entropy Change PerspectiveZhezheng Hao, Hong Wang, Haoyang Liu, Jian Luo 等ACL 2026 · 被引用 42 次
- Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM ReasoningShenzhi Wang, Le Yu, Chang Gao, Chujie Zheng 等NeurIPS 2025 · 被引用 592 次
- The Surprising Effectiveness of Negative Reinforcement in LLM ReasoningXinyu Zhu, Mengzhou Xia, Zhepei Wei, Wei-Lin Chen 等NeurIPS 2025 · 被引用 177 次
- Trust, But Verify: A Self-Verification Approach to Reinforcement Learning with Verifiable RewardsXiaoyuan Liu, Tian Liang, Zhiwei He, Jiahao Xu 等NeurIPS 2025 · 被引用 46 次
- Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?Zhiqi Chen, Rui Lu, Andrew Zhao, Zhaokai Wang 等NeurIPS 2025 · 被引用 1,109 次
