On the Generalization Gap in Self-Evolving Language Model Reasoning
Zhenting Qi, Susanna Maria Baby, Stefanie Baby, Kan Yuan, Andrew Tomkins, Tu Vu, Da-Cheng Juan, Cyrus Rashtchian
Abstract
Recent work suggests that large language models (LLMs) can improve through self-evolution (SE), using supervision signals generated by the model itself. In this work, we ask: under a strict closed-loop setup, where the SE algorithm has access only to an unlabeled prompt set and a base model, how close can internally generated supervision come to oracle-supervised training? We analyze four representative strategies in a unified offline self-evolution framework, including single-round verification, multi-turn revision with feedback, iterative training, and curriculum learning. Our primary experiments use Knights and Knaves (KK) logical reasoning tasks, which provide deterministic solutions, controlled difficulty levels, and a clean testbed for easy-to-hard generalization. We first show that SE consistently improves over the base model, but plateaus after excessive training compute is invested, and eventually still leaves a non trivial gap to oracle supervision. We find that multi-turn critic-revision with large models could reach strong self-evolution performance, where Gemma 12B nearly matches oracle-supervised training. Beyond KK, we also evaluate SE on real-world reasoning benchmarks, where gains are also modest. Overall, our results characterize when closed-loop SE can help, and show how internally generated supervision remains insufficient under this minimal formulation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 125c8a36-e5bc-41b7-9b12-b5e4edec9cf8Builds on27
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan et al.NeurIPS 2023 · 4,972 citations
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards et al.ICLR 2024 · 3,045 citations
- Absolute Zero: Reinforced Self-play Reasoning with Zero DataAndrew Zhao, Yiran Wu, Tong Wu, Quentin Xu et al.NeurIPS 2025 · 361 citations
- Rubrics as Rewards: Reinforcement Learning Beyond Verifiable DomainsAnisha Gunjal, Anthony Wang, Elaine Lau, Vaskar Nath et al.ICLR 2026 · 340 citations
Related papers
- Large Language Models Can Self-ImproveJiaxin Huang, Shixiang Gu, Le Hou, Yuexin Wu et al.EMNLP 2023 · 184 citations
- A Task-centric Theory for Iterative Self-Improvement with Easy-to-Hard CurriculaChenruo Liu, Yijun Dong, Yiqiu Shen, Qi LeiICML 2026
- Reward-Guided Prompt Evolving in Reinforcement Learning for LLMsZiyu Ye, Rishabh Agarwal, Tianqi Liu, Rishabh Joshi et al.ICML 2025
- Evoke: Evoking Critical Thinking Abilities in LLMs via Reviewer-Author Prompt EditingXinyu Hu, Pengfei Tang, Simiao Zuo, Zihan Wang et al.ICLR 2024 · 14 citations
- DEBATE, TRAIN, EVOLVE: Self-Evolution of Language Model ReasoningGaurav Srivastava, Zhenyu Bi, Meng Lu, Xuan WangEMNLP 2025 · 1 citation
