GUI-G1: Understanding R1-Zero-Like Training for Visual Grounding in GUI Agents
Yuqi Zhou, Sunhao Dai, Shuai Wang, Kaiwen Zhou, Qinglin Jia, Jun Xu
摘要
Recent Graphical User Interface (GUI) agents replicate the R1-Zero paradigm, coupling online Reinforcement Learning (RL) with explicit chain-of-thought reasoning prior to object grounding and thereby achieving substantial performance gains. In this paper, we first conduct extensive analysis experiments of three key components of that training pipeline: input design, output evaluation, and policy update-each revealing distinct challenges arising from blindly applying general-purpose RL without adapting to GUI grounding tasks. Input design: Current templates encourage the model to generate chain-of-thought reasoning, but longer chains unexpectedly lead to worse grounding performance. Output evaluation: Reward functions based on hit signals or box area allow models to exploit box size, leading to reward hacking and poor localization quality. Policy update: Online RL tends to overfit easy examples due to biases in length and sample difficulty, leading to under-optimization on harder cases. To address these issues, we propose three targeted solutions. First, we adopt a Fast Thinking Template that encourages direct answer generation, reducing excessive reasoning during training. Second, we incorporate a box size constraint into the reward function to mitigate reward hacking. Third, we revise the RL objective by adjusting length normalization and adding a difficulty-aware scaling factor, enabling better optimization on hard samples. Our GUI-G1-3B, trained on 17K public samples with Qwen2.5-VL-3B-Instruct, achieves 90.3% accuracy on ScreenSpot and 37.1% on ScreenSpot-Pro. This surpasses all prior models of similar size and even outperforms the larger UI-TARS-7B, establishing a new state-of-the-art in GUI agent grounding. The project repository is available at https://github.com/Yuqi-Zhou/GUI-G1.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper21
- GTA1: GUI Test-time Scaling AgentYan Yang, Dongxu Li, Yutong Dai, Yuhao Yang 等ICLR 2026 · 被引用 109 次
- ScaleCUA: Scaling Open-Source Computer Use Agents with Cross-Platform DataZhaoyang Liu, Jingjing Xie, Zichen Ding, Zehao Li 等ICLR 2026 · 被引用 54 次
- GUI-G²: Gaussian Reward Modeling for GUI GroundingFei Tang, Zhangxuan Gu, Zhengxi Lu, Xuyang Liu 等AAAI 2026 · 被引用 48 次
- DiMo-GUI: Advancing Test-time Scaling in GUI Grounding via Modality-Aware Visual ReasoningHang Wu, Hongkai Chen, Yujun Cai, Chang Liu 等EMNLP 2025 · 被引用 23 次
- UI-Ins: Enhancing GUI Grounding with Multi-Perspective Instruction as ReasoningLiangyu Chen, Hanzhang Zhou, chenglin Cai, Jianan Zhang 等ICLR 2026 · 被引用 19 次
它引用的顶会 Paper10
- Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language ModelsWenxuan Huang, Bohan Jia, Shaosheng Cao, Zheyu Ye 等ICLR 2026 · 被引用 670 次
- UI-R1: Enhancing Efficient Action Prediction of GUI Agents by Reinforcement LearningZhengxi Lu, Yuxiang Chai, Yaxuan Guo, Xi Yin 等AAAI 2026 · 被引用 103 次
- SeeClick: Harnessing GUI Grounding for Advanced Visual GUI AgentsKanzhi Cheng, Qiushi Sun, Yougang Chu, Fangzhi Xu 等ACL 2024 · 被引用 33 次
- ScreenSpot-Pro: GUI Grounding for Professional High-Resolution Computer UseKaixin Li, Ziyang Meng, Hongzhan Lin, Ziyang Luo 等ACM MM 2025 · 被引用 24 次
- Back to Basics: Revisiting REINFORCE-Style Optimization for Learning from Human Feedback in LLMsArash Ahmadian, Chris Cremer, Matthias Gallé, Marzieh Fadaee 等ACL 2024 · 被引用 20 次
相关 Paper
- GuirlVG: Incentivize GUI Visual Grounding via Empirical Exploration on Reinforcement LearningWeitai Kang, Bin Lei, Gaowen Liu, Caiwen Ding 等ICLR 2026 · 被引用 6 次
- GUI-Eyes: Tool-Augmented Perception for Visual Grounding in GUI AgentsChen Chen, Jiawei Shao, Dakuan Lu, Haoyi Hu 等AAAI 2026 · 被引用 5 次
- DRS-GUI: Dynamic Region Search for Training-Free GUI GroundingYichao Liu, Huawen Shen, Liu Yu, Shiyu Liu 等CVPR 2026 · 被引用 3 次
- FDC-Ground: Improving GRPO for GUI Grounding via Exponential Rewards and Fact-Aligned PruningXiangjian Zeng, Wenjing Li, Qingqiang Wu, Liang ZhangAAAI 2026 · 被引用 1 次
- SE-GUI: Enhancing Visual Grounding for GUI Agents via Self-Evolutionary Reinforcement LearningXinbin Yuan, Jian Jun Zhang, Kaixin Li, Zhuoxuan Cai 等NeurIPS 2025 · 被引用 4 次
