SeRL: Self-play Reinforcement Learning for Large Language Models with Limited Data
Wenkai Fang, Shunyu Liu, Yang Zhou, Kongcheng Zhang, Tongya Zheng, Kaixuan Chen, Mingli Song, Dacheng Tao
Abstract
Recent advances have demonstrated the effectiveness of Reinforcement Learning (RL) in improving the reasoning capabilities of Large Language Models (LLMs). However, existing works inevitably rely on high-quality instructions and verifiable rewards for effective training, both of which are often difficult to obtain in specialized domains. In this paper, we propose Self-play Reinforcement Learning (SeRL) to bootstrap LLM training with limited initial data. Specifically, SeRL comprises two complementary modules: self-instruction and self-rewarding. The former module generates additional instructions based on the available data at each training step, employing robust online filtering strategies to ensure instruction quality, diversity, and difficulty. The latter module introduces a simple yet effective majority-voting mechanism to estimate response rewards for additional instructions, eliminating the need for external annotations. Finally, SeRL performs conventional RL based on the generated data, facilitating iterative self-play learning. Extensive experiments on various reasoning benchmarks and across different LLM backbones demonstrate that the proposed SeRL yields results superior to its counterparts and achieves performance on par with those obtained by high-quality data with verifiable rewards. Our code is available at https://github.com/wantbook-book/SeRL.
However, in specialized domains such as clinical diagnostics [44,58,60] and aerospace engineering [5,36,75], collecting such high-quality data and verifiable supervision remains a significant challenge. This difficulty stems from the fact that acquiring human-labeled data in these areas is particularly labor-Corresponding author.
39th Conference on Neural Information Processing Systems (NeurIPS 2025).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 85fabc26-ad4a-4453-8799-c739b97a12c5Cited by top-tier papers15
- R-Zero: Self-Evolving Reasoning LLM from Zero DataChengsong Huang, Wenhao Yu, Xiaoyang Wang, Hongming Zhang et al.ICLR 2026 · 220 citations
- Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM ReasoningKongcheng Zhang, Qi Yao, Shunyu Liu, Yingjie Wang et al.NeurIPS 2025 · 45 citations
- Search Self-Play: Pushing the Frontier of Agent Capability without SupervisionHongliang Lu, Yuhang Wen, Pengyu Cheng, Ruijin Ding et al.ICLR 2026 · 35 citations
- How Far Can Unsupervised RLVR Scale LLM Training?Bingxiang He, Yuxin Zuo, Zeyuan Liu, Shangziqi Zhao et al.ICLR 2026 · 31 citations
- Co-rewarding: Stable Self-supervised RL for Eliciting Reasoning in Large Language ModelsZizhuo Zhang, Jianing Zhu, Xinmu Ge, Zihua Zhao et al.ICLR 2026 · 16 citations
Builds on25
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards et al.ICLR 2024 · 3,045 citations
- WizardLM: Empowering Large Pre-Trained Language Models to Follow Complex InstructionsCan Xu, Qingfeng Sun, Kai Zheng, Xiubo Geng et al.ICLR 2024 · 1,206 citations
- STaR: Bootstrapping Reasoning With ReasoningEric Zelikman, Yuhuai Wu, Jesse Mu, Noah D. GoodmanNeurIPS 2022 · 1,126 citations
- Self-Rewarding Language ModelsWeizhe Yuan, Richard Yuanzhe Pang, Kyunghyun Cho, Xian Li et al.ICML 2024 · 569 citations
- Self-Instruct: Aligning Language Models with Self-Generated InstructionsYizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu et al.ACL 2023 · 540 citations
Related papers
- Supervised Reinforcement Learning: From Expert Trajectories to Step-wise ReasoningYihe Deng, I-Hung Hsu, Jun Yan, Zifeng Wang et al.ICLR 2026 · 11 citations
- Learning from Synthetic Data Improves Multi-hop ReasoningAnmol Kabra, Yilun Yin, Albert Gong, Kamilė Stankevičiūtė et al.ICLR 2026 · 6 citations
- Is PRM Necessary? Problem-Solving RL Implicitly Induces PRM Capability in LLMsZhangyin Feng, Qianglong Chen, Ning Lu, Yongqian Li et al.NeurIPS 2025 · 16 citations
- Tailored Primitive Initialization is the Secret Key to Reinforcement LearningYihang Yao, Guangtao Zeng, Raina Wu, Yang Zhang et al.ACL 2026 · 1 citation
- General-Reasoner: Advancing LLM Reasoning Across All DomainsXueguang Ma, Qian Liu, Dongfu Jiang, Ge Zhang et al.NeurIPS 2025 · 153 citations
