Strategist: Self-improvement of LLM Decision Making via Bi-Level Tree Search
Jonathan Light, Min Cai, Weiqin Chen, Guanzhi Wang, Xiusi Chen, Wei Cheng, Yisong Yue, Ziniu Hu
摘要
Traditional reinforcement learning and planning typically requires vast amounts of data and training to develop effective policies. In contrast, large language models (LLMs) exhibit strong generalization and zero-shot capabilities, but struggle with tasks that require detailed planning and decision-making in complex action spaces. We introduce STRATEGIST, a novel approach that integrates the strengths of both methods. Our approach leverages LLMs to search and update high-level strategies (as text), which are then refined and executed by low-level Monte Carlo Tree Search (MCTS). STRATEGIST is a generalizable framework to optimize the strategy through population-based self-play simulations without the need for any training data. We demonstrate the effectiveness of STRATEGIST in learning optimal strategies for competitive, multi-turn games with partial information, including Game of Pure Strategy (GOPS) and multi-agent, hidden-identity discussion games like The Resistance: Avalon. Our results show that agents equipped with STRATE-GIST outperform those trained with traditional RL methods, other LLM-based skill acquisition techniques, pre-existing LLM agents across both game environments and achieves comparable performance against human players.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Scaling Inference-Time Computation via Opponent Simulation: Enabling Online Strategic Adaptation in Repeated NegotiationXiangyu Liu, Di Wang, Zhe Feng, Aranyak MehtaICML 2026 · 被引用 2 次
- The Stackelberg Speaker: Optimizing Persuasive Communication in Social Deduction GamesZheng Zhang, Deheng Ye, Peilin Zhao, Hao WangACL 2026
- Escaping Whack-a-Mole: Optimizing Documentation as Repo-Specific Playbooks for Coding AgentsYutong Cheng, Haifeng Chen, Wenchao Yu, Xujiang Zhao 等ICML 2026
- Cardiverse: Harnessing LLMs for Novel Card Game PrototypingDanrui Li, Sen Zhang, Samuel S. Sohn, Kaidong Hu 等EMNLP 2025
它引用的顶会 Paper15
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Toolformer: Language Models Can Teach Themselves to Use ToolsTimo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu 等NeurIPS 2023 · 被引用 5,989 次
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan 等NeurIPS 2023 · 被引用 5,828 次
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran 等NeurIPS 2023 · 被引用 5,068 次
相关 Paper
- Strat-Reasoner: Reinforcing Strategic Reasoning of LLMs in Multi-Agent GamesYidong He, Yutao Lai, Pengxu Yang, Jiarui Gan 等ICML 2026
- Strategic Planning: A Top-Down Approach to Option GenerationMax Ruiz Luyten, Antonin Berthon, Mihaela van der SchaarICML 2025
- Language Agents with Reinforcement Learning for Strategic Play in the Werewolf GameZelai Xu, Chao Yu, Fei Fang, Yu Wang 等ICML 2024 · 被引用 145 次
- Code World Models for General Game PlayingWolfgang Lehrach, Daniel Hennes, Miguel Lazaro-Gredilla, Xinghua Lou 等ICLR 2026 · 被引用 27 次
- MARSHAL: Incentivizing Multi-Agent Reasoning via Self-Play with Strategic LLMsHuining Yuan, Zelai Xu, Zheyue Tan, Xiangmin Yi 等ICLR 2026 · 被引用 15 次
