Cultivating Gaming Sense for Yourself: Making VLMs Gaming Experts
Wenxuan Lu, Jiangyang He, Zhanqiu Zhang, Steven Y. Guo, Tianning Zang
Abstract
Developing agents capable of fluid gameplay in first/third-person games without API access remains a critical challenge in Artificial General Intelligence (AGI). Recent efforts leverage Vision Language Models (VLMs) as direct controllers, frequently pausing the game to analyze screens and plan action through language reasoning. However, this inefficient paradigm fundamentally restricts agents to basic and nonfluent interactions: relying on isolated VLM reasoning for each action makes it impossible to handle tasks requiring high reactivity (e.g., FPS shooting) or dynamic adaptability (e.g., ACT combat). To handle this, we propose a paradigm shift in gameplay agent design: instead of directly controlling gameplay, VLM develops specialized execution modules tailored for tasks like shooting and combat. These modules handle real-time game interactions, elevating VLM to a high-level developer. Building upon this paradigm, we introduce GameSense, a gameplay agent framework where VLM develops task-specific game sense modules by observing task execution and leveraging vision tools and neural network training pipelines. These modules encapsulate action-feedback logic, ranging from direct action rules to neural network-based decisions. Experiments demonstrate that our framework is the first to achieve fluent gameplay in diverse genres, including ACT, FPS, and Flappy Bird, setting a new benchmark for game-playing agents. The code is available at https://github.com/ipsss2/GameSense .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e58f744b-2f26-4eae-93c5-bc44df842034Builds on2
Related papers
- GameplayQA: A Benchmarking Framework for Decision-Dense POV-Synced Multi-Video Understanding of 3D Virtual AgentsYunzhe Wang, Runhui Xu, Kexin Zheng, Tianyi Zhang et al.ACL 2026 · 2 citations
- CombatVLA: An Efficient Vision-Language-Action Model for Combat Tasks in 3D Action Role-Playing GamesPeng Chen, Pi Bu, Yingyao Wang, Xinyi Wang et al.ICCV 2025 · 1 citation
- Agentic Jigsaw Interaction Learning for Enhancing Visual Perception and Reasoning in Vision-Language ModelsYu Zeng, Wenxuan Huang, Shiting Huang, Xikun Bao et al.ICLR 2026 · 11 citations
- Real-Time Reasoning Agents in Evolving EnvironmentsYule Wen, Yixin Ye, Yanzhe Zhang, Diyi Yang et al.ICLR 2026 · 10 citations
- De-fine: Decomposing and Refining Visual Programs with Auto-FeedbackMinghe Gao, Juncheng Li, Hao Fei, Liang Pang et al.ACM MM 2024 · 1 citation
