Mastering Board Games by External and Internal Planning with Language Models
John Schultz, Jakub Adámek, Matej Jusup, Marc Lanctot, Michael Kaisers, Sarah Perrin, Daniel Hennes, Jeremy Shar, Cannada A. Lewis, Anian Ruoss, Tom Zahavy, Petar Velickovic
Abstract
Advancing planning and reasoning capabilities of Large Language Models (LLMs) is one of the key prerequisites towards unlocking their potential for performing reliably in complex and impactful domains. In this paper, we aim to demonstrate this across board games (Chess, Fischer Random / Chess960, Connect Four, and Hex), and we show that search-based planning can yield significant improvements in LLM game-playing strength. We introduce, compare and contrast two major approaches: In external search, the model guides Monte Carlo Tree Search (MCTS) rollouts and evaluations without calls to an external game engine, and in internal search, the model is trained to generate in-context a linearized tree of search and a resulting final choice. Both build on a language model pre-trained on relevant domain knowledge, reliably capturing the transition and value functions in the respective environments, with minimal hallucinations. We evaluate our LLM search implementations against game-specific state-of-the-art engines, showcasing substantial improvements in strength over the base model, and reaching Grandmaster-level performance in chess while operating closer to the human search budget. Our proposed approach, combining search with domain knowledge, is not specific to board games, hinting at more general future applications. * Equal contribution ** Equal senior authorship † Research conducted during an internship at Google 1 Google Deep-Mind 2 ETH Zürich 3 Google.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 235ff67f-d135-432a-83b5-593f15410b82Cited by top-tier papers6
- Code World Models for General Game PlayingWolfgang Lehrach, Daniel Hennes, Miguel Lazaro-Gredilla, Xinghua Lou et al.ICLR 2026 · 27 citations
- Look-ahead Reasoning with a Learned Model in Imperfect Information GamesOndrej Kubícek, Viliam LisýICLR 2026 · 3 citations
- Mixing Expert Knowledge: Bring Human Thoughts Back To the Game of GoYichuan Ma, Linyang Li, Yongkang Chen, Peiji Li et al.NeurIPS 2025 · 2 citations
- Out-of-Distribution Evaluation of Rule-Based and Strategic Reasoning in Chess TransformersAnna Mészáros, Patrik Reizinger, Ferenc HuszárICML 2026
- Cardiverse: Harnessing LLMs for Novel Card Game PrototypingDanrui Li, Sen Zhang, Samuel S. Sohn, Kaidong Hu et al.EMNLP 2025
Builds on44
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran et al.NeurIPS 2023 · 5,068 citations
- Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question AnsweringPan Lu, Swaroop Mishra, Tanglin Xia, Liang Qiu et al.NeurIPS 2022 · 2,727 citations
- Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought PromptingMiles Turpin, Julian Michael, Ethan Perez, Samuel R. BowmanNeurIPS 2023 · 1,792 citations
- Graph of Thoughts: Solving Elaborate Problems with Large Language ModelsMaciej Besta, Nils Blach, Ales Kubicek, Robert Gerstenberger et al.AAAI 2024 · 1,292 citations
Related papers
- Implicit Search via Discrete Diffusion: A Study on ChessJiacheng Ye, Zhenyu Wu, Jiahui Gao, Zhiyong Wu et al.ICLR 2025
- DSG-MCTS: A Dynamic Strategy-Guided Monte Carlo Tree Search for Diversified Reasoning in Large Language ModelsRui Ha, Chaozhuo Li, Rui Pu, Litian Zhang et al.EMNLP 2025
- Monte Carlo Planning with Large Language Model for Text-Based Game AgentsZijing Shi, Meng Fang, Ling ChenICLR 2025
- Strategist: Self-improvement of LLM Decision Making via Bi-Level Tree SearchJonathan Light, Min Cai, Weiqin Chen, Guanzhi Wang et al.ICLR 2025
- Toward Self-Improvement of LLMs via Imagination, Searching, and CriticizingYe Tian, Baolin Peng, Linfeng Song, Lifeng Jin et al.NeurIPS 2024 · 162 citations
