ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search
Dan Zhang, Sining Zhoubian, Ziniu Hu, Yisong Yue, Yuxiao Dong, Jie Tang
摘要
Recent methodologies in LLM self-training mostly rely on LLM generating responses and filtering those with correct output answers as training data. This approach often yields a low-quality fine-tuning training set (e.g., incorrect plans or intermediate reasoning). In this paper, we develop a reinforced self-training approach, called ReST-MCTS*, based on integrating process reward guidance with tree search MCTS* for collecting higher-quality reasoning traces as well as per-step value to train policy and reward models. ReST-MCTS* circumvents the per-step manual annotation typically used to train process rewards by tree-search-based reinforcement learning: Given oracle final correct answers, ReST-MCTS* is able to infer the correct process rewards by estimating the probability this step can help lead to the correct answer. These inferred rewards serve dual purposes: they act as value targets for further refining the process reward model and also facilitate the selection of high-quality traces for policy model self-training. We first show that the tree-search policy in ReST-MCTS* achieves higher accuracy compared with prior LLM reasoning baselines such as Best-of-N and Tree-of-Thought, within the same search budget. We then show that by using traces searched by this tree-search policy as training data, we can continuously enhance the three language models for multiple iterations, and outperform other self-training algorithms such as ReST and Self-Rewarding LM. We release all code at https://github.com/THUDM/ReST-MCTS.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper157
- Mulberry: Empowering MLLM with o1-like Reasoning and Reflection via Collective Monte Carlo Tree SearchHuanjin Yao, Jiaxing Huang, Wenhao Wu, Jingyi Zhang 等NeurIPS 2025 · 被引用 147 次
- Fast Best-of-N Decoding via Speculative RejectionHanshi Sun, Momin Haider, Ruiqi Zhang, Huitao Yang 等NeurIPS 2024 · 被引用 144 次
- Darwin Gödel Machine: Open-Ended Evolution of Self-Improving AgentsJenny Zhang, Shengran Hu, Cong Lu, Robert Tjarko Lange 等ICLR 2026 · 被引用 101 次
- Segment Policy Optimization: Effective Segment-Level Credit Assignment in RL for Large Language ModelsYiran Guo, Lijie Xu, Ji Liu, Dan Ye 等NeurIPS 2025 · 被引用 75 次
- Tree Search for LLM Agent Reinforcement LearningYuxiang Ji, Ziyu Ma, Yong Wang, Guanhua Chen 等ICLR 2026 · 被引用 71 次
它引用的顶会 Paper25
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran 等NeurIPS 2023 · 被引用 5,068 次
相关 Paper
- Learning to Better Search with Language Models via Guided Reinforced Self-TrainingSeungyong Moon, Bumsoo Park, Hyun Oh SongNeurIPS 2025 · 被引用 3 次
- TreeRL: LLM Reinforcement Learning with On-Policy Tree SearchZhenyu Hou, Ziniu Hu, Yujiang Li, Rui Lu 等ACL 2025
- K-STaR: Knowledge-Aware Self-Taught ReasonerGuozheng Li, Xinyu ZhangAAAI 2026
- Language Models can Self-Improve at State-Value Estimation for Better SearchEthan Mendes, Alan RitterNeurIPS 2025 · 被引用 5 次
- Re-ReST: Reflection-Reinforced Self-Training for Language AgentsZi-Yi Dou, Cheng-Fu Yang, Xueqing Wu, Kai-Wei Chang 等EMNLP 2024 · 被引用 2 次
