OCNR: Stabilizing Self-Play by Mitigating Iteration-Collapse With One-Class Novelty Rewards
Seungyoo Lee, Giung Nam, Hyungi Lee, Juho Lee
Abstract
Training large language models via self-play often suffers from a persistent iteration-collapse, where performance initially improves but subsequently regresses as training iterations increase. We analyze this phenomenon as arising from cross-iteration degeneration, where the task-generation distribution becomes increasingly confined to a narrow subset of familiar (seen) problems, weakening the effective learning signal and destabilizing training. To address this issue, we propose a plug-in approach that augments existing self-play pipelines with a one-class novelty reward. A Seen Detector trained on a historical buffer of previously used training problems identifies in-support instances and discourages redundant generation by the questioner, thereby steering exploration toward under-explored yet learnable regions. Experimental results show that the proposed method mitigates iteration-collapse during iterative training and yields consistent improvements.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5578aead-62c8-4604-a5e1-44f53ada3976Builds on13
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards et al.ICLR 2024 · 3,045 citations
- Solving Quantitative Reasoning Problems with Language ModelsAitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer et al.NeurIPS 2022 · 2,039 citations
- Self-Play Fine-Tuning Converts Weak Language Models to Strong Language ModelsZixiang Chen, Yihe Deng, Huizhuo Yuan, Kaixuan Ji et al.ICML 2024 · 527 citations
- Absolute Zero: Reinforced Self-play Reasoning with Zero DataAndrew Zhao, Yiran Wu, Tong Wu, Quentin Xu et al.NeurIPS 2025 · 361 citations
- TTRL: Test-Time Reinforcement LearningYuxin Zuo, Kaiyan Zhang, Li Sheng, Shang Qu et al.NeurIPS 2025 · 249 citations
Related papers
- Propose, Solve, Verify: Self-Play Through Formal VerificationAlex Wilf, Pranjal Aggarwal, Bryan Parno, Daniel Fried et al.ICML 2026 · 6 citations
- A Task-centric Theory for Iterative Self-Improvement with Easy-to-Hard CurriculaChenruo Liu, Yijun Dong, Yiqiu Shen, Qi LeiICML 2026
- Self-Improving Transformers Overcome Easy-to-Hard and Length Generalization ChallengesNayoung Lee, Ziyang Cai, Avi Schwarzschild, Kangwook Lee et al.ICML 2025
- Mind the Gap: Examining the Self-Improvement Capabilities of Large Language ModelsYuda Song, Hanlin Zhang, Carson Eisenach, Sham M. Kakade et al.ICLR 2025 · 3 citations
- Beyond Pass@ 1: Self-Play with Variational Problem Synthesis Sustains RLVRXiao Liang, Zhong-Zhi Li, Yeyun Gong, Yelong Shen et al.ICLR 2026 · 57 citations
