GAMEBoT: Transparent Assessment of LLM Reasoning in Games
Wenye Lin, Jonathan Roberts, Yunhan Yang, Samuel Albanie, Zongqing Lu, Kai Han
Abstract
Large Language Models (LLMs) are increasingly deployed in real-world applications that demand complex reasoning. To track progress, robust benchmarks are required to evaluate their capabilities beyond superficial pattern recognition. However, current LLM reasoning benchmarks often face challenges such as insufficient interpretability, performance saturation or data contamination. To address these challenges, we introduce GAMEBOT (GAME Battle of Tactics), a gaming arena designed for rigorous and transparent assessment of LLM reasoning capabilities. GAMEBOT decomposes complex reasoning in games into predefined modular subproblems. This decomposition allows us to design a suite of Chain-of-Thought (CoT) prompts that leverage domain knowledge to guide LLMs in addressing these subproblems before action selection. Furthermore, we develop a suite of rule-based algorithms to generate ground truth for these subproblems, enabling rigorous validation of the LLMs' intermediate reasoning steps. This approach facilitates evaluation of both the quality of final actions and the accuracy of the underlying reasoning process. GAMEBOT also naturally alleviates the risk of data contamination through dynamic games and head-to-head LLM competitions. We benchmark 17 prominent LLMs across eight games, encompassing various strategic abilities and game characteristics. Our results suggest that GAMEBOT presents a significant challenge, even when LLMs are provided with detailed CoT prompts. Project page: https://visual-ai.github.io/gamebot LLMs CoT Prompts <Role Setting> You are an expert player in … <Game Rules> <Inputs and denotations> You will receive the current game state denoted … <Output> Provide your chosen move. Before making a decision, articulate your internal thinking process. Your performance will be assessed on both the intermediate thinking results and the final decision. Follow the thinking process: * [Intermediate Thinking Results 1] * [Intermediate Thinking Results 2] … [Chosen Move] Competitive Game Environments Othello Pong Surround Checkers TicTacToe Connect4 Texas hold'em Negotiate v2 Games Game Properties Representative Abilites Avg. Turns Action Space State Space Type Information Simul. Zero-sum
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0f4dba27-637b-41db-a534-3203f9a2fa58Cited by top-tier papers3
- How Far Are LLMs from Professional Poker Players? Revisiting Game-Theoretic Reasoning with Agentic Tool UseMinhua Lin, Enyan Dai, Hui Liu, Xianfeng Tang et al.ICLR 2026 · 9 citations
- Evaluating Language Models' Evaluations of GamesKatherine M. Collins, Cedegao E. Zhang, Graham Todd, Lance Ying et al.ICLR 2026 · 5 citations
- MeepleLM: A Virtual Playtester Simulating Diverse Subjective ExperiencesZizhen Li, Chuanhao Li, Yibin Wang, Jianwen Sun et al.ACL 2026 · 1 citation
Builds on9
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo et al.NeurIPS 2022 · 8,168 citations
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan et al.NeurIPS 2023 · 5,828 citations
- Chatbot Arena: An Open Platform for Evaluating LLMs by Human PreferenceWei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos et al.ICML 2024 · 1,212 citations
- AgentBench: Evaluating LLMs as AgentsXiao Liu, Hao Yu, Hanchen Zhang, Yifan Xu et al.ICLR 2024 · 748 citations
Related papers
- GameArena: Evaluating LLM Reasoning through Live Computer GamesLanxiang Hu, Qiyu Li, Anze Xie, Nan Jiang et al.ICLR 2025
- Evaluating the Inductive Abilities of Large Language Models: Why Chain-of-Thought Reasoning Sometimes Hurts More Than HelpsHaibo Jin, Peiyan Zhang, Man Luo, Haohan WangNeurIPS 2025 · 1 citation
- GTBench: Uncovering the Strategic Reasoning Capabilities of LLMs via Game-Theoretic EvaluationsJinhao Duan, Renming Zhang, James Diffenderfer, Bhavya Kailkhura et al.NeurIPS 2024 · 79 citations
- Are Large Vision Language Models Good Game Players?Xinyu Wang, Bohan Zhuang, Qi WuICLR 2025
- GuessArena: Guess Who I Am? A Self-Adaptive Framework for Evaluating LLMs in Domain-Specific Knowledge and ReasoningQingchen Yu, Zifan Zheng, Ding Chen, Simin Niu et al.ACL 2025 · 5 citations
