ReGenesis: LLMs can Grow into Reasoning Generalists via Self-Improvement
Xiangyu Peng, Congying Xia, Xinyi Yang, Caiming Xiong, Chien-Sheng Wu, Chen Xing
Abstract
Post-training Large Language Models (LLMs) with explicit reasoning trajectories can enhance their reasoning abilities. However, acquiring such high-quality trajectory data typically demands meticulous supervision from humans or superior models, which can be either expensive or license-constrained. In this paper, we explore how far an LLM can improve its reasoning by self-synthesizing reasoning paths as training data without any additional supervision. Existing self-synthesizing methods, such as STaR, suffer from poor generalization to out-of-domain (OOD) reasoning tasks. We hypothesize it is due to that their self-synthesized reasoning paths are too task-specific, lacking general task-agnostic reasoning guidance. To address this, we propose Reasoning Generalist via Self-Improvement (ReGenesis 1 ) , a method to self-synthesize reasoning paths as post-training data by progressing from abstract to concrete. More specifically, ReGenesis self-synthesizes reasoning paths by converting general reasoning guidelines into task-specific ones, generating reasoning structures, and subsequently transforming these structures into reasoning paths, without the need for human-designed task-specific examples used in existing methods. We show that ReGenesis achieves superior performance on all in-domain and OOD settings tested compared to existing methods. For six OOD tasks specifically, while previous methods exhibited an average performance decrease of approximately 4.6% after post training, ReGenesis delivers around 6.1% performance improvement. We also conduct in-depth analysis of our framework and show ReGenesis is effective across various LLMs and design choices.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1e39d873-fab5-41b6-9f3f-efa278fafc13Cited by top-tier papers9
- Learning to Reason via Mixture-of-Thought for Logical ReasoningTong Zheng, Lichang Chen, Simeng Han, R. Thomas McCoy et al.ICLR 2026 · 23 citations
- Diversity-Enhanced Reasoning for Subjective QuestionsYumeng Wang, Zhiyuan Fan, Jiayu Liu, Jen-Tse Huang et al.ICLR 2026 · 13 citations
- RuleReasoner: Reinforced Rule-based Reasoning via Domain-aware Dynamic SamplingYang Liu, Jiaqi Li, Zilong ZhengICLR 2026 · 8 citations
- AdaSTaR: Adaptive Data Sampling for Training Self-Taught ReasonersReiss Koh, Wonbeen Oh, Jaein Jang, Minhyung Lee et al.NeurIPS 2025 · 8 citations
- Better, Faster: Harnessing Self-Improvement in Large Reasoning ModelsQihuang Zhong, Liang Ding, Juhua Liu, Bo Du et al.ICML 2026 · 3 citations
Builds on16
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo et al.NeurIPS 2022 · 8,168 citations
- STaR: Bootstrapping Reasoning With ReasoningEric Zelikman, Yuhuai Wu, Jesse Mu, Noah D. GoodmanNeurIPS 2022 · 1,126 citations
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
- CRITIC: Large Language Models Can Self-Correct with Tool-Interactive CritiquingZhibin Gou, Zhihong Shao, Yeyun Gong, Yelong Shen et al.ICLR 2024 · 699 citations
Related papers
- Genius: A Generalizable and Purely Unsupervised Self-Training Framework For Advanced ReasoningFangzhi Xu, Hang Yan, Chang Ma, Haiteng Zhao et al.ACL 2025 · 22 citations
- K-STaR: Knowledge-Aware Self-Taught ReasonerGuozheng Li, Xinyu ZhangAAAI 2026
- Learn to Think: Improving Multimodal Reasoning through Vision-Aware Self-Improvement TrainingQihuang Zhong, Liang Ding, Wenjie Xuan, Juhua Liu et al.ICML 2026 · 1 citation
- Tool-Star: Empowering Multi-Tool Collaborative Web Agent via Reinforcement LearningGuanting Dong, Yifei Chen, Xiaoxi Li, Jiajie Jin et al.SIGIR 2026 · 1 citation
- Large Language and Reasoning Models are Shallow Disjunctive ReasonersIrtaza Khalid, Amir Masoud Nourollah, Steven SchockaertACL 2025
