Seeing Without Understanding: Disentangling Perception, Reasoning, and Simulation in VLM Gameplay
Dingyang Jin, Jiawei He, Calvin Lo, Steven Hu, RYAN RAD
摘要
While Vision-Language Models (VLMs) excel on static visual benchmarks, they consistently underperform in game-based reasoning, yet existing evaluations conflate failures in perception, rule comprehension, and reasoning. We propose a two-stage diagnostic framework that decomposes VLM performance into testable components: controlled perception tests isolating visual encoding, and a diagnostic matrix with a six-level rule complexity ladder evaluated in both explicit verification and predictive simulation modes. Experimenting with six state-of-the-art VLMs reveals three failure patterns: (1) coordinated spatial drift, where off-by-one localization errors among adjacent pieces share the same shift direction at - the rate expected under spatial independence; (2) perception-reasoning dissociation, where models correctly verify board states but fail to apply rules—at complex constraint levels, perception remains relatively stable while reasoning accuracy plummets, with even the best-performing model capped at ; and (3) a simulation gap, with performance dropping by up to points when predicting future states versus verifying observed outcomes. These limitations persist across model scales and are not resolved by scaling, text-only input, or structured prompting. Code and data are available at https://github.com/chillibeaver/PRS-Diag.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper6
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic ManipulationHao Shi, Bin Xie, Yingfei Liu, Lin Sun 等ICLR 2026 · 被引用 227 次
- VisuRiddles: Fine-grained Perception is a Primary Bottleneck for Multimodal Large Language Models in Abstract Visual ReasoningHao Yan, Xingchen Liu, Hao Wang, Zhenbiao Cao 等ICLR 2026 · 被引用 7 次
- VS-Bench: Evaluating VLMs for Strategic Abilities in Multi-Agent EnvironmentsZelai Xu, Zhexuan Xu, Xiangmin Yi, Huining Yuan 等CVPR 2026 · 被引用 3 次
相关 Paper
- Are Large Vision Language Models Good Game Players?Xinyu Wang, Bohan Zhuang, Qi WuICLR 2025
- Diagnosing the Compositional Knowledge of Vision Language Models from a Game-Theoretic ViewJin Wang, Shichao Dong, Yapeng Zhu, Kelu Yao 等ICML 2024 · 被引用 6 次
- Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language ModelsZesen Lyu, Dandan Zhang, Wei Ye, Fangdi Li 等EMNLP 2025
- VisRes Bench: On Evaluating the Visual Reasoning Capabilities of VLMsBrigitta Malagurski Törtei, Yasser Dahou, Ngoc Dung Huynh, Wamiq Reyaz Para 等CVPR 2026 · 被引用 3 次
- LLMs as Rules Oracles: Exploring Real-World Multimodal Reasoning in Tabletop Strategy Game EnvironmentsJoseph Peper, Sai Krishna Gandra, Yunxiang Zhang, Vaibhav Chennareddy 等ICLR 2026
