RethinkMCTS: Refining Erroneous Thoughts in Monte Carlo Tree Search for Code Generation
Qingyao Li, Wei Xia, Xinyi Dai, Kounianhua Du, Weiwen Liu, Yasheng Wang, Ruiming Tang, Yong Yu, Weinan Zhang
Abstract
Tree search methods have demonstrated impressive performance in code generation. Previous methods combine tree search with reflection that summarizes past mistakes to achieve iterative improvement. However, these methods face significant challenges. First, they search directly within the code language space, neglecting the underlying reasoning process critical for effective code generation. Second, reflection-based approaches merely accumulate historical errors in memory without providing correct reasoning pathways, making it difficult for subsequent search iterations to identify optimal solutions, resulting in decreased search quality. In this work, we propose RETHINKMCTS, a framework that systematically explores and refines the reasoning process for code generation. Specifically, we employ MCTS to search for thoughts before code generation and integrate MCTS with a refinement mechanism called rethink, which incorporates fine-grained code execution feedback to refine erroneous thoughts during the search. It ensures the search path aligns with better reasoning, improving overall search quality. Through extensive experiments, we demonstrate that RETHINKMCTS outperforms previous search-based and feedback-enhanced code generation baselines 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 47bb1034-e7e3-4a46-a694-a7fb218284f9Cited by top-tier papers10
- From Large to Small: Transferring CUDA Optimization Expertise via Reasoning GraphJunfeng Gong, Zhiyi Wei, Junying Chen, Cheng Liu et al.ICLR 2026 · 10 citations
- Let's Revise Step-by-Step: A Unified Local Search Framework for Code Generation with LLMsZhiyi Lyu, Jianguo Huang, Yanchen Deng, Steven Hoi et al.NeurIPS 2025 · 7 citations
- ATGen: Adversarial Reinforcement Learning for Test Case GenerationQingyao Li, Xinyi Dai, Weiwen Liu, Xiangyang Li et al.ICLR 2026 · 4 citations
- MCTS-Refined CoT: High-Quality Fine-Tuning Data for LLM-Based Repository Issue ResolutionYibo Wang, Zhihao Peng, Ying Wang, Zhao Wei et al.ASE 2025 · 4 citations
- ReSo: A Reward-driven Self-organizing LLM-based Multi-Agent System for Reasoning TasksHeng Zhou, Hejia Geng, Xiangyuan Xue, Li Kang et al.EMNLP 2025 · 4 citations
Builds on12
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan et al.NeurIPS 2023 · 5,828 citations
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran et al.NeurIPS 2023 · 5,068 citations
- WizardCoder: Empowering Code Large Language Models with Evol-InstructZiyang Luo, Can Xu, Pu Zhao, Qingfeng Sun et al.ICLR 2024 · 945 citations
- Language Agent Tree Search Unifies Reasoning, Acting, and Planning in Language ModelsAndy Zhou, Kai Yan, Michal Shlapentokh-Rothman, Haohan Wang et al.ICML 2024 · 443 citations
- AlphaZero-Like Tree-Search can Guide Large Language Model Decoding and TrainingZiyu Wan, Xidong Feng, Muning Wen, Stephen Marcus McAleer et al.ICML 2024 · 325 citations
Related papers
- RPM-MCTS: Knowledge-Retrieval as Process Reward Model with Monte Carlo Tree Search for Code GenerationYuanyuan Lin, Xiangyu Ouyang, Teng Zhang, Kaixin SuiAAAI 2026 · 1 citation
- MARS²: Scaling Multi-Agent Tree Search via Reinforcement Learning for Code GenerationPengfei Li, Shijie Wang, Fangyuan Li, Yikun Fu et al.ACL 2026
- ReflectionCoder: Learning from Reflection Sequence for Enhanced One-off Code GenerationHouxing Ren, Mingjie Zhan, Zhongyuan Wu, Aojun Zhou et al.ACL 2025 · 12 citations
- SeDev: Structured Semantic Exploration for LLM-Driven Code GenerationRonghui Yang, Jie Liu, Jiajie Zeng, Jiexin Wang et al.ACL 2026
- SE-Agent: Self-Evolution Trajectory Optimization in Multi-Step Reasoning with LLM-Based AgentsYifu Guo, Jiaye Lin, Huacan Wang, Yuzhen Han et al.NeurIPS 2025 · 73 citations
