Superficial Self-Improved Reasoners Benefit from Model Merging
Xiangchi Yuan, Chunhui Zhang, Zheyuan Liu, Dachuan Shi, Leyan Pan, Soroush Vosoughi, Wenke Lee
Abstract
As scaled language models (LMs) approach human-level reasoning capabilities, selfimprovement emerges as a solution to synthesizing high-quality data corpus. While previous research has identified model collapse as a risk in self-improvement, where model outputs become increasingly deterministic, we discover a more fundamental challenge: the superficial self-improved reasoners phenomenon. In particular, our analysis reveals that even when LMs show improved in-domain (ID) reasoning accuracy, they actually compromise their generalized reasoning capabilities on out-of-domain (OOD) tasks due to memorization rather than genuine learning. Through a systematic investigation of LM architecture, we discover that during self-improvement, LM weight updates are concentrated in less reasoning-critical layers, leading to superficial learning. To address this, we propose Iterative Model Merging (IMM), a method that strategically combines weights from original and self-improved models to preserve generalization while incorporating genuine reasoning improvements. Our approach effectively mitigates both LM collapse and superficial learning, moving towards more stable self-improving systems. Code is available 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers12
- SwiReasoning: Switch-Thinking in Latent and Explicit for Pareto-Superior Reasoning LLMsDachuan Shi, Abedelkadir Asi, Keying Li, Xiangchi Yuan et al.ICLR 2026 · 22 citations
- Behavior Knowledge Merge in Reinforced Agentic ModelsXiangchi Yuan, Dachuan Shi, Chunhui Zhang, Zheyuan Liu et al.ACL 2026 · 7 citations
- Growing Through Experience: Scaling Episodic Grounding in Language ModelsChunhui Zhang, Sirui Wang, Zhongyu Ouyang, Xiangchi Yuan et al.ACL 2025 · 6 citations
- Reasoning over Precedents Alongside Statutes: Case-Augmented Deliberative Alignment for LLM SafetyCan Jin, Rui Wu, Tong Che, Qixin Zhang et al.ACL 2026 · 3 citations
- Stabilizing Self-Consuming Diffusion Models with Latent Space FilteringZhongteng Cai, Yaxuan Wang, Yang Liu, Xueru ZhangAAAI 2026 · 2 citations
Builds on39
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo et al.NeurIPS 2022 · 8,168 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran et al.NeurIPS 2023 · 5,068 citations
- STaR: Bootstrapping Reasoning With ReasoningEric Zelikman, Yuhuai Wu, Jesse Mu, Noah D. GoodmanNeurIPS 2022 · 1,126 citations
Related papers
- MindMerger: Efficiently Boosting LLM Reasoning in non-English LanguagesZixian Huang, Wenhao Zhu, Gong Cheng, Lei Li et al.NeurIPS 2024 · 31 citations
- Progress or Regress? Self-Improvement Reversal in Post-trainingTing Wu, Xuefeng Li, Pengfei LiuICLR 2025
- Can MLLMs Absorb Math Reasoning Abilities from LLMs as Free Lunch?Yijie Hu, Zihao Zhou, Kaizhu Huang, Xiaowei Huang et al.NeurIPS 2025 · 3 citations
- Self-MoE: Towards Compositional Large Language Models with Self-Specialized ExpertsJunmo Kang, Leonid Karlinsky, Hongyin Luo, Zhen Wang et al.ICLR 2025
- Beyond Layer-Wise Merging: Chain-of-Merging for Vision-Language ModelsXinyu Zhang, Yuxuan Dong, Lingling Zhang, Chengyou Jia et al.CVPR 2026
