SE-GUI: Enhancing Visual Grounding for GUI Agents via Self-Evolutionary Reinforcement Learning
Xinbin Yuan, Jian Jun Zhang, Kaixin Li, Zhuoxuan Cai, Lujian Yao, Jie Chen, Enguang Wang, Qibin Hou, Jinwei Chen, Peng-Tao Jiang, Bo Li
Abstract
Graphical User Interface (GUI) agents have made substantial strides in understanding and executing user instructions across diverse platforms. Yet, grounding these instructions to precise interface elements remains challenging-especially in complex, high-resolution, professional environments. Traditional supervised fine-tuning (SFT) methods often require large volumes of diverse data and exhibit weak generalization. To overcome these limitations, we introduce a reinforcement learning (RL)-based framework that incorporates three core strategies: (1) seed data curation to ensure high-quality training samples, (2) a dense policy gradient that provides continuous feedback based on prediction accuracy, and (3) a self-evolutionary reinforcement finetuning mechanism that iteratively refines the model using attention maps. With only 3k training samples, our 7B-parameter model achieves state-of-the-art results among similarly sized models on three grounding benchmarks. Notably, it attains 47.3% accuracy on the ScreenSpot-Pro dataset-outperforming much larger models, such as UI-TARS-72B, by a margin of 24.2%. These findings underscore the effectiveness of RL-based approaches in enhancing GUI agent performance, particularly in high-resolution, complex environments.
- This work was done during Xinbin Yuan's internship at vivo. ‡ Project lead.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5af7b2e5-6528-4dc4-ba6c-d4c4542b0191Cited by top-tier papers1
Ask how each one uses itBuilds on13
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language ModelsWenxuan Huang, Bohan Jia, Shaosheng Cao, Zheyu Ye et al.ICLR 2026 · 670 citations
- Visual-RFT: Visual Reinforcement Fine-TuningZiyu Liu, Zeyi Sun, Yuhang Zang, Xiaoyi Dong et al.ICCV 2025 · 563 citations
- Mobile-Agent-v2: Mobile Device Operation Assistant with Effective Navigation via Multi-Agent CollaborationJunyang Wang, Haiyang Xu, Haitao Jia, Xi Zhang et al.NeurIPS 2024 · 245 citations
- UI-R1: Enhancing Efficient Action Prediction of GUI Agents by Reinforcement LearningZhengxi Lu, Yuxiang Chai, Yaxuan Guo, Xi Yin et al.AAAI 2026 · 103 citations
Related papers
- GuirlVG: Incentivize GUI Visual Grounding via Empirical Exploration on Reinforcement LearningWeitai Kang, Bin Lei, Gaowen Liu, Caiwen Ding et al.ICLR 2026 · 6 citations
- UI-Ins: Enhancing GUI Grounding with Multi-Perspective Instruction as ReasoningLiangyu Chen, Hanzhang Zhou, chenglin Cai, Jianan Zhang et al.ICLR 2026 · 19 citations
- GUI-G1: Understanding R1-Zero-Like Training for Visual Grounding in GUI AgentsYuqi Zhou, Sunhao Dai, Shuai Wang, Kaiwen Zhou et al.NeurIPS 2025 · 73 citations
- SeeClick: Harnessing GUI Grounding for Advanced Visual GUI AgentsKanzhi Cheng, Qiushi Sun, Yougang Chu, Fangzhi Xu et al.ACL 2024 · 33 citations
- GUI-Bee: Align GUI Action Grounding to Novel Environments via Autonomous ExplorationYue Fan, Handong Zhao, Ruiyi Zhang, Yu Shen et al.EMNLP 2025 · 15 citations
