GUI-Rise: Structured Reasoning and History Summarization for GUI Navigation
Tao Liu, Chongyu Wang, Rongjie Li, Yingchen Yu, Xuming He, Song Bai
摘要
While Multimodal Large Language Models (MLLMs) have advanced GUI navigation agents, current approaches face limitations in cross-domain generalization and effective history utilization. We present a reasoning-enhanced framework that systematically integrates structured reasoning, action prediction, and history summarization. The structured reasoning component generates coherent Chain-of-Thought analyses combining progress estimation and decision reasoning, which inform both immediate action predictions and compact history summaries for future steps. Based on this framework, we train a GUI agent, GUI-Rise, through supervised fine-tuning on pseudo-labeled trajectories and reinforcement learning with Group Relative Policy Optimization (GRPO). This framework employs specialized rewards, including a history-aware objective, directly linking summary quality to subsequent action performance. Comprehensive evaluations on standard benchmarks demonstrate state-of-the-art results under identical training data conditions, with particularly strong performance in out-of-domain scenarios. These findings validate our framework's ability to maintain robust reasoning and generalization across diverse GUI navigation tasks. Code is available at https://leon022.github.io/GUI-Rise.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- UI-Copilot: Advancing Long-Horizon GUI Automation via Tool-Integrated Policy OptimizationZhengxi Lu, Fei Tang, Guangyi Liu, Jin Ma 等ACL 2026 · 被引用 2 次
- Experience-driven Multi-turn Reinforcement Learning for GUI AgentsZhengxi Lu, Jiabo Ye, Fei Tang, Yongliang Shen 等ACL 2026
它引用的顶会 Paper12
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan 等NeurIPS 2025 · 被引用 2,828 次
- GPT-4V(ision) is a Generalist Web Agent, if GroundedBoyuan Zheng, Boyu Gou, Jihyung Kil, Huan Sun 等ICML 2024 · 被引用 496 次
- AppAgent: Multimodal Agents as Smartphone UsersChi Zhang, Zhao Yang, Jiaxuan Liu, Yanda Li 等CHI 2025 · 被引用 57 次
- SeeClick: Harnessing GUI Grounding for Advanced Visual GUI AgentsKanzhi Cheng, Qiushi Sun, Yougang Chu, Fangzhi Xu 等ACL 2024 · 被引用 33 次
- MobileGPT: Augmenting LLM with Human-like App Memory for Mobile Task AutomationSunjae Lee, Junyoung Choi, Jungjae Lee, Munim Hasan Wasi 等MobiCom 2024 · 被引用 21 次
相关 Paper
- History-Aware Reasoning for GUI AgentsZiwei Wang, Leyang Yang, Xiaoxuan Tang, Sheng Zhou 等AAAI 2026
- UI-R1: Enhancing Efficient Action Prediction of GUI Agents by Reinforcement LearningZhengxi Lu, Yuxiang Chai, Yaxuan Guo, Xi Yin 等AAAI 2026 · 被引用 103 次
- Enhancing GUI Agent with Uncertainty-Aware Self-Trained EvaluatorGongwei Chen, Lirong Jie, Lexiao Zou, Weili Guan 等NeurIPS 2025 · 被引用 4 次
- Look Before You Leap: A GUI-Critic-R1 Model for Pre-Operative Error Diagnosis in GUI AutomationYuyang Wanyan, Xi Zhang, Haiyang Xu, Haowei Liu 等NeurIPS 2025 · 被引用 26 次
- Plan Then Action: High-Level Planning Guidance Reinforcement Learning for LLM ReasoningZhihao Dou, Qinjian Zhao, Zhongwei Wan, Zhang Dinggen 等ICML 2026 · 被引用 24 次
