R-Zero: Self-Evolving Reasoning LLM from Zero Data
Chengsong Huang, Wenhao Yu, Xiaoyang Wang, Hongming Zhang, Zongxia Li, Ruosen Li, Jiaxin Huang, Haitao Mi, Dong Yu
Abstract
Self-evolving Large Language Models (LLMs) offer a scalable path toward superintelligence by autonomously generating, refining, and learning from their own experiences. However, existing methods for training such models still rely heavily on vast human curated tasks and labels, typically via fine-tuning or reinforcement learning, which poses a fundamental bottleneck to advancing AI systems toward capabilities beyond human intelligence. To overcome this limitation, we introduce R-Zero, a fully autonomous framework that generates its own training data from scratch. Starting from a single base LLM, R-Zero initializes two independent models with distinct roles -a Challenger and a Solver. These models are optimized separately and co-evolve through interaction: the Challenger is rewarded for proposing tasks near the edge of the Solver's capability, and the Solver is rewarded for solving increasingly challenging tasks posed by the Challenger. This process yields a targeted, self-improving curriculum without any pre-existing tasks and labels. Empirically, R-Zero substantially improves reasoning capability across different backbone LLMs, e.g., boosting the Qwen3-4B-Base by +6.49 on math reasoning benchmarks, and +7.54 on general-domain reasoning benchmarks. Code: https://github.com/Chengsong-Huang/R-Zero . Figure 1: (Left): R-Zero employs a co-evolutionary loop between Challenger and Solver. (Right): R-Zero achieves strong benchmark gains without any pre-existing tasks or human labels.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f3da9eaa-3593-488c-895c-722c89595343Cited by top-tier papers18
- VideoHallu: Evaluating and Mitigating Multi-modal Hallucinations on Synthetic Video UnderstandingZongxia Li, Xiyang Wu, Guangyao Shi, Yubin Qin et al.NeurIPS 2025 · 38 citations
- UniCorn: Towards Self-Improving Unified Multimodal Models through Self-Generated SupervisionZhen Fang, Ruiyan Han, XinYu Sun, Yuchen Ma et al.ACL 2026 · 19 citations
- Teaching Models to Teach Themselves: Reasoning at the Edge of LearnabilityShobhita Sundaram, John Quan, Ariel Kwiatkowski, Kartik Ahuja et al.ICML 2026 · 16 citations
- Agentic Proposing: Enhancing Large language Model Reasoning via Compositional Skill SynthesisZhengbo Jiao, Shaobo Wang, Zifan Zhang, Xuan Ren et al.ICML 2026 · 10 citations
- CoVerRL: Breaking the Consensus Trap in Label-Free Reasoning via Generator-Verifier Co-EvolutionTeng Pan, Yuchen Yan, Zixuan Wang, Ruiqing Zhang et al.ACL 2026 · 5 citations
Builds on27
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Solving Quantitative Reasoning Problems with Language ModelsAitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer et al.NeurIPS 2022 · 2,039 citations
- Self-Play Fine-Tuning Converts Weak Language Models to Strong Language ModelsZixiang Chen, Yihe Deng, Huizhuo Yuan, Kaixuan Ji et al.ICML 2024 · 527 citations
- Absolute Zero: Reinforced Self-play Reasoning with Zero DataAndrew Zhao, Yiran Wu, Tong Wu, Quentin Xu et al.NeurIPS 2025 · 361 citations
- TTRL: Test-Time Reinforcement LearningYuxin Zuo, Kaiyan Zhang, Li Sheng, Shang Qu et al.NeurIPS 2025 · 249 citations
Related papers
- General-Reasoner: Advancing LLM Reasoning Across All DomainsXueguang Ma, Qian Liu, Dongfu Jiang, Ge Zhang et al.NeurIPS 2025 · 153 citations
- Think before Recommendation: Autonomous Reasoning-enhanced RecommenderXiaoyu Kong, Junguang Jiang, Bin Liu, Ziru Xu et al.NeurIPS 2025 · 17 citations
- D²Evo: Dual Difficulty-Aware Self-Evolution for Data-Efficient Reinforcement LearningRu Zhang, Renda Li, Ziyu Ma, Weijie Qiu et al.ICML 2026
- SPIRAL: Self-Play on Zero-Sum Games Incentivizes Reasoning via Multi-Agent Multi-Turn Reinforcement LearningBo Liu, Simon Yu, Zichen Liu, Leon Guertler et al.ICLR 2026 · 88 citations
- PretrainZero: Reinforcement Active Learning on Pretraining DataXingrun Xing, Zhiyuan Fan, Jie Lou, Guoqi Li et al.ICML 2026
