GameVerse: Can Vision-Language Models Learn from Video-based Reflection?
Kuan Zhang, Dongchen Liu, Qiyue Zhao, Jinkun Hou, Xinran Zhang, Qinlei Xie, Miao Liu, Yiming Li
摘要
Human gameplay is a visually grounded interaction loop in which players act, reflect on failures, and watch tutorials to refine strategies. Can Vision-Language Models (VLMs) also learn from video-based reflection? We present GameVerse , a comprehensive video game benchmark that enables a reflective visual interaction loop . Moving beyond traditional fire-and-forget evaluations, it uses a novel reflect-and-retry paradigm to assess how VLMs internalize visual experience and improve policies. To facilitate systematic and scalable evaluation, we also introduce a cognitive hierarchical taxonomy spanning 15 globally popular games, dual action space for both semantic and GUI control, and milestone evaluation using advanced VLMs to quantify progress. Our experiments show that VLMs benefit from video-based reflection in varied settings, and perform best by combining failure trajectories and expert tutorials—a training-free analogue to reinforcement learning (RL) plus supervised fine-tuning (SFT).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper11
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan 等NeurIPS 2023 · 被引用 5,828 次
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan 等NeurIPS 2023 · 被引用 4,972 次
- GROOT: Learning to Follow Instructions by Watching Gameplay VideosShaofei Cai, Bowei Zhang, Zihao Wang, Xiaojian Ma 等ICLR 2024 · 被引用 43 次
- ScreenSpot-Pro: GUI Grounding for Professional High-Resolution Computer UseKaixin Li, Ziyang Meng, Hongzhan Lin, Ziyang Luo 等ACM MM 2025 · 被引用 24 次
- Real-Time Reasoning Agents in Evolving EnvironmentsYule Wen, Yixin Ye, Yanzhe Zhang, Diyi Yang 等ICLR 2026 · 被引用 10 次
相关 Paper
- Game-RL: Synthesizing Multimodal Verifiable Game Data to Boost VLMs' General ReasoningJingqi Tong, Jixin Tang, Hangcheng Li, Yurong Mou 等ICLR 2026 · 被引用 21 次
- Game Ground Bench: Probing the Limits of LVLMs in Complex Semantic Grounding Across Game UniversesZhangyang Qi, Jinsong Li, Hongjian Wu, Jiaqi Wang 等AAAI 2026
- ProgressLM: Towards Progress Reasoning in Vision-Language ModelsJianshu Zhang, Chengxuan Qian, Haosen Sun, Haoran Lu 等ACL 2026 · 被引用 7 次
- SRPO: Enhancing Multimodal LLM Reasoning via Reflection-Aware Reinforcement LearningZhongwei Wan, Zhihao Dou, Che Liu, Yu Zhang 等NeurIPS 2025 · 被引用 63 次
- VERSE: Verification-based Self-Play for Code InstructionsHao Jiang, Qi Liu, Rui Li, Yuze Zhao 等AAAI 2025 · 被引用 3 次
