FlashAdventure: A Benchmark for GUI Agents Solving Full Story Arcs in Diverse Adventure Games
Jaewoo Ahn, Junseo Kim, Heeseung Yun, Jaehyeon Son, Dongmin Park, Jaewoong Cho, Gunhee Kim
Abstract
GUI agents powered by LLMs show promise in interacting with diverse digital environments. Among these, video games offer a valuable testbed due to their varied interfaces, with adventure games posing additional challenges through complex, narrative-driven interactions. Existing game benchmarks, however, lack diversity and rarely evaluate agents on completing entire storylines. To address this, we introduce FlashAdventure, a benchmark of 34 Flashbased adventure games designed to test full story arc completion and tackle the observationbehavior gap: the challenge of remembering and acting on earlier gameplay information. We also propose CUA-as-a-Judge, an automated gameplay evaluator, and COAST, an agentic framework leveraging long-term clue memory to better plan and solve sequential tasks. Experiments show current GUI agents struggle with full story arcs, while COAST improves milestone completion by bridging the observationbehavior gap. Nonetheless, a marked discrepancy between humans and best-performing agents warrants continued research efforts to narrow this divide. * Equal contribution. †Work done during an internship at KRAFTON.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b21d4627-7f26-48b1-b880-5a60c6f3bc9dCited by top-tier papers2
- Orak: A Foundational Benchmark for Training and Evaluating LLM Agents on Diverse Video GamesDongmin Park, Minkyu Kim, Beongjun Choi, Junhyuck Kim et al.ICLR 2026 · 30 citations
- GameVerse: Can Vision-Language Models Learn from Video-based Reflection?Kuan Zhang, Dongchen Liu, Qiyue Zhao, Jinkun Hou et al.ICML 2026 · 2 citations
Builds on22
- Generative Agents: Interactive Simulacra of Human BehaviorJoon Sung Park, Joseph C. O'Brien, Carrie Jun Cai, Meredith Ringel Morris et al.UIST 2023 · 1,882 citations
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
- Interactive Fiction Games: A Colossal AdventureMatthew J. Hausknecht, Prithviraj Ammanabrolu, Marc-Alexandre Côté, Xingdi YuanAAAI 2020 · 242 citations
- Large Language Models Play StarCraft II: Benchmarks and A Chain of Summarization ApproachWeiyu Ma, Qirui Mi, Yongcheng Zeng, Xue Yan et al.NeurIPS 2024 · 122 citations
- SmartPlay : A Benchmark for LLMs as Intelligent AgentsYue Wu, Xuan Tang, Tom M. Mitchell, Yuanzhi LiICLR 2024 · 121 citations
Related papers
- PlayCoder: Making LLM-Generated GUI Code PlayableZhiyuan Peng, Wei Tao, Xin Yin, Chenhao Ying et al.FSE 2026
- Benchmarking Agent Memory in Interdependent Multi-Session Agentic TasksZexue He, Yu Wang, Churan Zhi, Yuanzhe Hu et al.ICML 2026 · 2 citations
- EscapeBench: Towards Advancing Creative Intelligence of Language Model AgentsCheng Qian, Peixuan Han, Qinyu Luo, Bingxiang He et al.ACL 2025
- UltraHorizon: Benchmarking LLM-Agent Capabilities in Ultra Long-Horizon ScenariosHaotian Luo, Huaisong Zhang, Xuelin Zhang, Haoyu Wang et al.ICML 2026 · 21 citations
- GUI-CEval: A Hierarchical and Comprehensive Chinese Benchmark for Mobile GUI AgentsYang Li, Yuchen Liu, Haoyu Lu, Zhiqiang Xia et al.CVPR 2026 · 3 citations
