GameplayQA: A Benchmarking Framework for Decision-Dense POV-Synced Multi-Video Understanding of 3D Virtual Agents
Yunzhe Wang, Runhui Xu, Kexin Zheng, Tianyi Zhang, Jayavibhav Niranjan Kogundi, Soham Hans, Volkan Ustun
Abstract
Multimodal LLMs are increasingly deployed as perceptual backbones for autonomous agents in 3D environments, from robotics to virtual worlds. These applications require agents to perceive rapid state changes, attribute actions to the correct entities, and reason about concurrent multi-agent behaviors from a first-person perspective, capabilities that existing benchmarks do not adequately evaluate. We introduce GameplayQA, a framework for evaluating agentic-centric perception and reasoning through video understanding. Specifically, we densely annotate multiplayer 3D gameplay videos at 1.22 labels/second, with time-synced, concurrent captions of states, actions, and events structured around a triadic system of Self, Other Agents, and the World, a natural decomposition for multi-agent environments. From these annotations, we refined 2.4K diagnostic QA pairs organized into three levels of cognitive complexity, accompanied by a structured distractor taxonomy that enables fine-grained analysis of where models hallucinate. Evaluation of frontier MLLMs reveals a substantial gap from human performance, with common failures in temporal and cross-video grounding, agent-role attribution, and handling the decision density of the game. We hope GameplayQA stimulates future research at the intersection of embodied AI, agentic perception, and world modeling.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on8
- EgoHumans: An Egocentric 3D Multi-Human BenchmarkRawal Khirodkar, Aayush Bansal, Lingni Ma, Richard A. Newcombe et al.ICCV 2023 · 59 citations
- OpenEQA: Embodied Question Answering in the Era of Foundation ModelsArjun Majumdar, Anurag Ajay, Xiaohan Zhang, Pranav Putta et al.CVPR 2024 · 44 citations
- WebVoyager: Building an End-to-End Web Agent with Large Multimodal ModelsHongliang He, Wenlin Yao, Kaixin Ma, Wenhao Yu et al.ACL 2024 · 30 citations
- GlitchBench: Can Large Multimodal Models Detect Video Game Glitches?Mohammad Reza Taesiri, Tianjun Feng, Cor-Paul Bezemer, Anh NguyenCVPR 2024 · 7 citations
- EGOILLUSION: Benchmarking Hallucinations in Egocentric Video UnderstandingAshish Seth, Utkarsh Tyagi, Ramaneswaran Selvakumar, Nishit Anand et al.EMNLP 2025 · 1 citation
Related papers
- Out of Sight, Not Out of Context? Egocentric Spatial Reasoning in VLMs Across Disjoint FramesSahithya Ravi, Gabriel Herbert Sarch, Vibhav Vineet, Andrew D. Wilson et al.EMNLP 2025 · 1 citation
- Embodied VideoAgent: Persistent Memory from Egocentric Videos and Embodied Sensors Enables Dynamic Scene UnderstandingYue Fan, Xiaojian Ma, Rongpeng Su, Jun Guo et al.ICCV 2025 · 2 citations
- Seeing, Listening, Remembering, and Reasoning: A Multimodal Agent with Long-Term MemoryLin Long, Yichen He, Wentao Ye, Yiyuan Pan et al.ICLR 2026 · 90 citations
- Are Large Vision Language Models Good Game Players?Xinyu Wang, Bohan Zhuang, Qi WuICLR 2025
- A Multi-Agent Perception-Action Alliance for Efficient Long Video ReasoningYichang Xu, Gaowen Liu, Ramana Rao Kompella, Tiansheng Huang et al.CVPR 2026 · 2 citations
