NaturalGAIA: A Verifiable Benchmark and Hierarchical Framework for Long-Horizon GUI Tasks
Zihan Zheng, Tianle Cui, Taoran Wang, Fengtao Wang, Jiahui Pan, Lewei He, Qianglong Chen
Abstract
Despite significant advances in LLM-driven GUI agents, the field remains constrained by the challenge of reconciling high-fidelity realism with verifiable evaluation accuracy. To address this, we introduce NaturalGAIA, a verifiable evaluation dataset grounded in real-world human GUI interaction intents. By decoupling logical causal pathways from linguistic narratives, it rigorously simulates natural human intent, characterized by cognitive non-linearity and contextual dependencies. Furthermore, we propose LightManus-Jarvis, a hierarchical collaborative framework where LightManus manages dynamic topological planning and context evolution, while Jarvis ensures execution precision via hybrid visual-structural perception. Experiments demonstrate that our approach achieves a Weighted Pathway Success Rate of 45.6%, significantly outperforming the state-of-the-art baseline (21.1%), while reducing token consumption by 75% and execution time by 76%. These results validate the efficacy of the macro-planning and micro-execution paradigm in handling complex naturalized tasks. Our code is publicly available at: https://github.com/KeLes-Coding/NatureGAIA.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on14
- Agentic Context Engineering: Evolving Contexts for Self-Improving Language ModelsQizheng Zhang, Changran Hu, Shubhangi Upasani, Boyuan Ma et al.ICLR 2026 · 374 citations
- PlanGenLLMs: A Modern Survey of LLM Planning CapabilitiesHui Wei, Zihao Zhang, Shenghua He, Tian Xia et al.ACL 2025 · 78 citations
- SeeClick: Harnessing GUI Grounding for Advanced Visual GUI AgentsKanzhi Cheng, Qiushi Sun, Yougang Chu, Fangzhi Xu et al.ACL 2024 · 33 citations
- WebVoyager: Building an End-to-End Web Agent with Large Multimodal ModelsHongliang He, Wenlin Yao, Kaixin Ma, Wenhao Yu et al.ACL 2024 · 30 citations
- RealWebAssist: A Benchmark for Long-Horizon Web Assistance with Real-World UsersSuyu Ye, Haojun Shi, Darren Shih, Hyokun Yun et al.AAAI 2026 · 17 citations
Related papers
- Aguvis: Unified Pure Vision Agents for Autonomous GUI InteractionYiheng Xu, Zekun Wang, Junli Wang, Dunjie Lu et al.ICML 2025
- NL Schedule: Evaluate Multitask Scheduling Capability of Large Language ModelsWenrui Liao, Weihong Du, Yi Li, Hongru Liang et al.ACL 2026
- Agent-SAMA: State-Aware Mobile AssistantLinqiang Guo, Wei Liu, Yi Wen Heng, Tse-Hsun (Peter) Chen et al.AAAI 2026 · 2 citations
- GUI-World: A Video Benchmark and Dataset for Multimodal GUI-oriented UnderstandingDongping Chen, Yue Huang, Siyuan Wu, Jingyu Tang et al.ICLR 2025 · 1 citation
- COLA: Collaborative Multi-Agent Framework with Dynamic Task Scheduling for GUI AutomationDi Zhao, Longhui Ma, Siwei Wang, Miao Wang et al.EMNLP 2025
