LongHorizonUI: A Unified Framework for Robust long-horizon Task Automation of GUI Agent
Bin Kang, Shaoguo Wen, Yifei Bi, Shunlong Wu, Xinbin Yuan, Rui Shao, Junle Wang, Zhuotao Tian
Abstract
While multimodal large language models (MLLMs) have shown promise in short-horizon GUI agents, their performance degrades significantly on longhorizon tasks involving complex, dynamic interfaces. To address this, we present LongHorizonUI, a framework designed to enhance the reliability and robustness of MLLM-based agents in extended interactive environments. Moreover, we establish a new long-horizon benchmark, named LongGUIBench, encompassing complex general applications and various gaming scenarios. Long-horizon tasks in this benchmark are defined as those requiring more than 15 steps, enabling thorough evaluation of long-horizon reasoning capabilities. Building upon this benchmark, we develop a Multimodal Enhanced Perceiver that integrates element detection and text recognition models, assigning unique indices to interface elements, thereby reinforcing state representation. Furthermore, we introduce a Deep-Reflection Decider, which employs a structured multi-level feedback-validation mechanism to support iterative reasoning and guarantee precise action execution along predictable trajectories. Building on the Deciders outputs, a Compensatory Action Executor continuously monitors execution progress; when degradation is detected, it applies targeted compensation operations or triggers a rollback procedure, thereby maintaining robustness throughout long-horizon tasks. Experiments show that LongHorizonUI substantially improves long-horizon performance on LongGUIBench, while remaining competitive on diverse public benchmarks. The code is publicly available at https://kane2kang.github.io/ LongHorizonUI/ .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5d2d2731-1948-447f-8064-0abbd93ca344Cited by top-tier papers1
Ask how each one uses itBuilds on31
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Language Models can Solve Computer TasksGeunwoo Kim, Pierre Baldi, Stephen McAleerNeurIPS 2023 · 539 citations
- Mobile-Agent-v2: Mobile Device Operation Assistant with Effective Navigation via Multi-Agent CollaborationJunyang Wang, Haiyang Xu, Haitao Jia, Xi Zhang et al.NeurIPS 2024 · 245 citations
- Multimodal Web Navigation with Instruction-Finetuned Foundation ModelsHiroki Furuta, Kuang-Huei Lee, Ofir Nachum, Yutaka Matsuo et al.ICLR 2024 · 160 citations
Related papers
- MobileUse: A Hierarchical Reflection-Driven GUI Agent for Autonomous Mobile OperationNing Li, Xiangmou Qu, Jiamu Zhou, Muning Wen et al.NeurIPS 2025 · 42 citations
- GUI-CEval: A Hierarchical and Comprehensive Chinese Benchmark for Mobile GUI AgentsYang Li, Yuchen Liu, Haoyu Lu, Zhiqiang Xia et al.CVPR 2026 · 3 citations
- Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal AgentsTianyi Men, Zhuoran Jin, Pengfei Cao, Yubo Chen et al.ACL 2025 · 13 citations
- GUI-Reflection: Empowering Multimodal GUI Models with Self-Reflection BehaviorPenghao Wu, Shengnan Ma, Bo Wang, Jiaheng Yu et al.NeurIPS 2025 · 20 citations
- PG-Agent: An Agent Powered by Page GraphWeizhi Chen, Ziwei Wang, Leyang Yang, Sheng Zhou et al.ACM MM 2025
