LLMs as Rules Oracles: Exploring Real-World Multimodal Reasoning in Tabletop Strategy Game Environments
Joseph Peper, Sai Krishna Gandra, Yunxiang Zhang, Vaibhav Chennareddy, Shloki Jha, Ali Payani, Lu Wang
Abstract
We introduce LUDOBENCH 1 , a multimodal reasoning benchmark that evaluates whether vision-enabled large language models (LMs) can acquire, integrate, and reason over heterogeneous game knowledge in mainstream analog tabletop games. Unlike prior works that emphasize deep strategic mastery, LUDOBENCH targets an initial reasoning challenge uninitiated gamers face: correctly comprehending a new tabletop strategy game for the first time. We examine whether, given a visual depiction of a tabletop scene and a corresponding ruleset, a model can correctly answer grounded questions about the pictured scenario. Concretely, LUDOBENCH tests three cumulative situated game-comprehension capabilities: (1) Environment Perception, (2) Heterogeneous Rules Integration, and (3) Short-horizon Optimization, to progressively stress-test the foundational reasoning required for real-world game comprehension. Evaluating frontier LMs on five diverse strategy games, we find that even the strongest models achieve only ∼76% accuracy on simple environment perception tasks and fall below 13% on situated multi-step comprehension puzzles that hobbyist gamers can routinely solve. Our extensive failure analysis and knowledge-ablation experiments reveal that models largely fail to comprehend rich cross-modal reference knowledge and are subsequently unable to apply this knowledge to messy and unfamiliar situated environments. Our findings highlight the many steps remaining for current methods to succeed on complex multimodal reasoning in the real world. Benchmark, leaderboard, and visualizer are available at https://huggingface.co/spaces/launch/LudoBench .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 96a390c2-a796-48f6-b7d2-acab935e56f3Builds on3
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Benchmarking the Spectrum of Agent CapabilitiesDanijar HafnerICLR 2022 · 193 citations
- Are Large Vision Language Models Good Game Players?Xinyu Wang, Bohan Zhuang, Qi WuICLR 2025
Related papers
- GlitchBench: Can Large Multimodal Models Detect Video Game Glitches?Mohammad Reza Taesiri, Tianjun Feng, Cor-Paul Bezemer, Anh NguyenCVPR 2024 · 7 citations
- VS-Bench: Evaluating VLMs for Strategic Abilities in Multi-Agent EnvironmentsZelai Xu, Zhexuan Xu, Xiangmin Yi, Huining Yuan et al.CVPR 2026 · 3 citations
- lmgame-Bench: How Good are LLMs at Playing Games?Lanxiang Hu, Mingjia Huo, Yuxuan Zhang, Haoyang Yu et al.ICLR 2026 · 47 citations
- BALROG: Benchmarking Agentic LLM and VLM Reasoning On GamesDavide Paglieri, Bartlomiej Cupial, Samuel Coward, Ulyana Piterbarg et al.ICLR 2025
- VisuLogic: A Benchmark for Evaluating Visual Reasoning in Multi-modal Large Language ModelsWeiye Xu, Jiahao Wang, Weiyun Wang, Zhe Chen et al.ICLR 2026 · 103 citations
