GuirlVG: Incentivize GUI Visual Grounding via Empirical Exploration on Reinforcement Learning
Weitai Kang, Bin Lei, Gaowen Liu, Caiwen Ding, Yan Yan
摘要
Graphical user interface visual grounding (GUI-VG)-a core capability for GUI agents-has primarily relied on supervised fine-tuning (SFT) of multimodal large language models (MLLMs), demanding extensive data curation and significant training costs. However, as MLLMs continue to advance and even cover GUI domains during pretraining, the necessity of exhaustive SFT post-training becomes increasingly questionable. Meanwhile, the recent successes of rule-based reinforcement fine-tuning (RFT) suggest a more efficient alternative. However, despite its promise, the optimal manner of RFT for GUI-VG remains unexplored. To bridge this gap, we introduce GuirlVG, a reinforcement learning-based GUI-VG method built on a systematic empirical study and a novel stabilization technique. Preliminarily, we find that naive application of RFT underperforms the SFT baseline, motivating a deeper exploration of RFT. First, we decompose RFT into its core components and analyze the optimal formulation of each. Second, as part of this exploration, we propose a novel Adversarial KL Factor that dynamically stabilizes training to mitigate reward over-optimization. Third, we further explore the training configurations of RFT to enhance the effectiveness. Extensive experiments show that GuirlVG, with only 5.2K training samples, outperforms SFT methods trained on over 10M samples, achieving a +7.7% improvement on ScreenSpot, a +17.2% improvement on ScreenSpotPro and 91.9% accuracy on ScreenSpotV2. RFT (Trivial) 79.2% 83.4% + Tune β + Adv. KL + Img. Res. Figure 1: Step-by-step exploration of GuirlVG. Starting from trivial RFT, we progressively add Soft Reward Function, In-Bbox reward with point prediction, β tuning, our Adversarial KL Factor, image resolution prompting, and extended training. With only 5.2K data, GuirlVG surpasses SFT methods trained on up to 13.58M data. Circle size reflects data scale used by each method. Preprint. Under review.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- GUI-Spotlight: Adaptive Iterative Focus Refinement for Enhanced GUI Visual GroundingBin Lei, Nuo Xu, Ali Payani, Mingyi Hong 等ICML 2026 · 被引用 8 次
- VGent: Visual Grounding via Modular Design for Disentangling Reasoning and PredictionWeitai Kang, Jason Kuen, Mengwei Ren, Zijun Wei 等CVPR 2026 · 被引用 5 次
- Measure Twice, Click Once: Co-evolving Proposer and Visual Critic via Reinforcement Learning for GUI GroundingWenkai Wang, Xiyun Li, Hongcan Guo, Wenhao Yu 等ACL 2026 · 被引用 1 次
它引用的顶会 Paper17
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language ModelsWenxuan Huang, Bohan Jia, Shaosheng Cao, Zheyu Ye 等ICLR 2026 · 被引用 670 次
- Visual-RFT: Visual Reinforcement Fine-TuningZiyu Liu, Zeyi Sun, Yuhang Zang, Xiaoyi Dong 等ICCV 2025 · 被引用 563 次
- TransVG: End-to-End Visual Grounding with TransformersJiajun Deng, Zhengyuan Yang, Tianlang Chen, Wengang Zhou 等ICCV 2021 · 被引用 468 次
相关 Paper
- UI-R1: Enhancing Efficient Action Prediction of GUI Agents by Reinforcement LearningZhengxi Lu, Yuxiang Chai, Yaxuan Guo, Xi Yin 等AAAI 2026 · 被引用 103 次
- SE-GUI: Enhancing Visual Grounding for GUI Agents via Self-Evolutionary Reinforcement LearningXinbin Yuan, Jian Jun Zhang, Kaixin Li, Zhuoxuan Cai 等NeurIPS 2025 · 被引用 4 次
- GUI-G1: Understanding R1-Zero-Like Training for Visual Grounding in GUI AgentsYuqi Zhou, Sunhao Dai, Shuai Wang, Kaiwen Zhou 等NeurIPS 2025 · 被引用 73 次
- GUI-Eyes: Tool-Augmented Perception for Visual Grounding in GUI AgentsChen Chen, Jiawei Shao, Dakuan Lu, Haoyi Hu 等AAAI 2026 · 被引用 5 次
- GUI-Shift: Enhancing VLM-Based GUI Agents through Self-supervised Reinforcement LearningLongxi Gao, Li Zhang, Pengzhi Gao, Wei Liu 等ICLR 2026 · 被引用 11 次
