HiconAgent: History Context-aware Policy Optimization for GUI Agents
Xurui Zhou, Gongwei Chen, Yuquan Xie, Zaijing Li, Kaiwen Zhou, Shuai Wang, Shuo Yang, Zhuotao Tian, Rui Shao
摘要
Graphical User Interface (GUI) agents require effective utilization of historical context to perform sequential navigation tasks. While incorporating past actions and observations can significantly improve decision-making, naively using full history leads to excessive computational overhead and potential distraction from irrelevant information. In this work, we introduce HiconAgent , a GUI agent trained with History Context-aware Policy Optimization (HCPO) for effective and efficient utilization of historical information. HCPO explicitly optimizes history usage in both sampling and policy updates by integrating two complementary components: (1) Dynamic Context Sampling (DCS) presents the agent with variable-length histories during sampling, enabling adaptive use of the most relevant historical context to improve sequential decision quality; (2) Anchor-guided History Compression (AHC) refines the policy update phase via a dual-branch optimization strategy, where the compressed branch drops history observations while keeping history actions as information flow anchors. The compressed and uncompressed branches are coupled through a history-enhanced alignment loss to enforce consistent history usage, achieving efficiency with minimal performance degradation. Extensive experiments on mainstream GUI navigation benchmarks demonstrate the strong performance of our model. Despite its smaller size, HiconAgent-3B outperforms GUI-R1-7B by +8.46 % grounding and +11.32 % step successful rate on GUI-Odyssey, while achieving comparable results on AndroidControl and AITW, with up to 2.47× computational speedup and 60% FLOPs reduction .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic ManipulationZaijing Li, Bing Hu, Rui Shao, Gongwei Chen 等CVPR 2026 · 被引用 23 次
- ConsisVLA-4D: Advancing Spatiotemporal Consistency in Efficient 3D-Perception and 4D-Reasoning for Robotic ManipulationWei Li, Jizhihui Liu, Yixing Li, Junwen Tong 等CVPR 2026 · 被引用 8 次
- HATS: Hardness-Aware Trajectory Synthesis for GUI AgentsRui Shao, Ruize Gao, Bin Xie, Yixing Li 等CVPR 2026 · 被引用 7 次
- UI-Copilot: Advancing Long-Horizon GUI Automation via Tool-Integrated Policy OptimizationZhengxi Lu, Fei Tang, Guangyi Liu, Jin Ma 等ACL 2026 · 被引用 2 次
它引用的顶会 Paper29
- Visual-RFT: Visual Reinforcement Fine-TuningZiyu Liu, Zeyi Sun, Yuhang Zang, Xiaoyi Dong 等ICCV 2025 · 被引用 563 次
- Optimus-1: Hybrid Multimodal Memory Empowered Agents Excel in Long-Horizon TasksZaijing Li, Yuquan Xie, Rui Shao, Gongwei Chen 等NeurIPS 2024 · 被引用 104 次
- UI-R1: Enhancing Efficient Action Prediction of GUI Agents by Reinforcement LearningZhengxi Lu, Yuxiang Chai, Yaxuan Guo, Xi Yin 等AAAI 2026 · 被引用 103 次
- CogVLA: Cognition-Aligned Vision-Language-Action Models via Instruction-Driven Routing & SparsificationWei Li, Renshan Zhang, Rui Shao, Jie He 等NeurIPS 2025 · 被引用 87 次
- MoME: Mixture of Multimodal Experts for Generalist Multimodal Large Language ModelsLeyang Shen, Gongwei Chen, Rui Shao, Weili Guan 等NeurIPS 2024 · 被引用 55 次
相关 Paper
- Less is More: Empowering GUI Agent with Context-Aware SimplificationGongwei Chen, Xurui Zhou, Rui Shao, Yibo Lyu 等ICCV 2025 · 被引用 2 次
- GUIOdyssey: A Comprehensive Dataset for Cross-App GUI Navigation on Mobile DevicesQuanfeng Lu, Wenqi Shao, Zitao Liu, Lingxiao Du 等ICCV 2025 · 被引用 7 次
- GUI-Rise: Structured Reasoning and History Summarization for GUI NavigationTao Liu, Chongyu Wang, Rongjie Li, Yingchen Yu 等NeurIPS 2025 · 被引用 4 次
- UI-Hawk: Unleashing the Screen Stream Understanding for Mobile GUI AgentsJiwen Zhang, Ya-Qi Yu, Minghui Liao, WenTao Li 等EMNLP 2025 · 被引用 1 次
- SE-GA: Memory-Augmented Self-Evolution for GUI Agentsshilong jin, Lanjun Wang, Zhuosheng ZhangICML 2026 · 被引用 1 次
