Improving Rationality in the Reasoning Process of Language Models through Self-playing Game
Pinzheng Wang, Juntao Li, Zecheng Tang, Haijia Gui, Min Zhang
Abstract
Large language models (LLMs) have demonstrated considerable reasoning abilities in various tasks such as mathematics and coding. However, recent studies indicate that even the best models lack true comprehension of their reasoning processes. In this paper, we explore how self-play can enhance the rationality of models in the reasoning process without supervision from humans or superior models. We design a Critic-Discernment Game (CDG) in which a prover first provides a solution to a given problem and is subsequently challenged by critiques of its solution. These critiques either aim to assist or mislead the prover. The objective of the prover is to maintain the correct answer when faced with misleading comments, while correcting errors in response to constructive feedback. Our experiments on tasks involving mathematical reasoning, stepwise error detection, self-correction, and long-chain reasoning demonstrate that CDG training can significantly improve the ability of well-aligned LLMs to comprehend their reasoning process. We have released our code here.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on11
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards et al.ICLR 2024 · 3,045 citations
- ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree SearchDan Zhang, Sining Zhoubian, Ziniu Hu, Yisong Yue et al.NeurIPS 2024 · 527 citations
- Self-Play Fine-Tuning Converts Weak Language Models to Strong Language ModelsZixiang Chen, Yihe Deng, Huizhuo Yuan, Kaixuan Ji et al.ICML 2024 · 527 citations
Related papers
- SPC: Evolving Self-Play Critic via Adversarial Games for LLM ReasoningJiaqi Chen, Bang Zhang, Ruotian Ma, Peisong Wang et al.NeurIPS 2025 · 43 citations
- Self-playing Adversarial Language Game Enhances LLM ReasoningPengyu Cheng, Tianhao Hu, Han Xu, Zhisong Zhang et al.NeurIPS 2024 · 120 citations
- S^3cMath: Spontaneous Step-Level Self-Correction Makes Large Language Models Better Mathematical ReasonersYuchen Yan, Jin Jiang, Yang Liu, Yixin Cao et al.AAAI 2025 · 19 citations
- Is PRM Necessary? Problem-Solving RL Implicitly Induces PRM Capability in LLMsZhangyin Feng, Qianglong Chen, Ning Lu, Yongqian Li et al.NeurIPS 2025 · 16 citations
- Toward Self-Improvement of LLMs via Imagination, Searching, and CriticizingYe Tian, Baolin Peng, Linfeng Song, Lifeng Jin et al.NeurIPS 2024 · 162 citations
