When is Tree Search Useful for LLM Planning? It Depends on the Discriminator
Ziru Chen, Michael White, Raymond J. Mooney, Ali Payani, Yu Su, Huan Sun
摘要
In this paper, we examine how large language models (LLMs) solve multi-step problems under a language agent framework with three components: a generator, a discriminator, and a planning method. We investigate the practical utility of two advanced planning methods, iterative correction and tree search. We present a comprehensive analysis of how discrimination accuracy affects the overall performance of agents when using these two methods or a simpler method, re-ranking. Experiments on two tasks, text-to-SQL parsing and mathematical reasoning, show that: (1) advanced planning methods demand discriminators with at least 90% accuracy to achieve significant improvements over re-ranking; (2) current LLMs' discrimination abilities have not met the needs of advanced planning methods to achieve such improvements; (3) with LLM-based discriminators, advanced planning methods may not adequately balance accuracy and efficiency. For example, compared to the other two methods, tree search is at least 10-20 times slower but leads to negligible performance gains, which hinders its real-world applications. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Deliberate Reasoning in Language Models as Structure-Aware Planning with an Accurate World ModelSiheng Xiong, Ali Payani, Yuan Yang, Faramarz FekriACL 2025 · 被引用 25 次
- Aristotle: Mastering Logical Reasoning with A Logic-Complete Decompose-Search-Resolve FrameworkJundong Xu, Hao Fei, Meng Luo, Qian Liu 等ACL 2025 · 被引用 11 次
- ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific DiscoveryZiru Chen, Shijie Chen, Yuting Ning, Qianheng Zhang 等ICLR 2025 · 被引用 6 次
- Language Models can Self-Improve at State-Value Estimation for Better SearchEthan Mendes, Alan RitterNeurIPS 2025 · 被引用 5 次
- Encoder of Thoughts: Enhancing Planning Ability in Language Agents Through Structural EmbeddingYuxiang Zhang, Jitao SangAAAI 2025 · 被引用 1 次
它引用的顶会 Paper19
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan 等NeurIPS 2023 · 被引用 5,828 次
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran 等NeurIPS 2023 · 被引用 5,068 次
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan 等NeurIPS 2023 · 被引用 4,972 次
相关 Paper
- Tree-Planner: Efficient Close-loop Task Planning with Large Language ModelsMengkang Hu, Yao Mu, Xinmiao Yu, Mingyu Ding 等ICLR 2024 · 被引用 57 次
- Can LLMs Fix Issues with Reasoning Models? Towards More Likely Models for AI PlanningTurgay Caglar, Sirine Belhaj, Tathagata Chakraborti, Michael Katz 等AAAI 2024 · 被引用 11 次
- Small LLMs Are Weak Tool Learners: A Multi-LLM AgentWeizhou Shen, Chenliang Li, Hongzhan Chen, Ming Yan 等EMNLP 2024 · 被引用 18 次
- Semantic Exploration with Adaptive Gating for Efficient Problem Solving with Language ModelsSungjae Lee, Hyejin Park, Jaechang Kim, Jungseul OkACL 2025
- Policy Guided Tree Search for Enhanced LLM ReasoningYang LiICML 2025
