GUI-Shift: Enhancing VLM-Based GUI Agents through Self-supervised Reinforcement Learning
Longxi Gao, Li Zhang, Pengzhi Gao, Wei Liu, Jian Luan, Mengwei Xu
摘要
Training effective Vision-Language Models (VLMs) for GUI agents typically depends on large-scale annotated datasets, whose collection is both labor-intensive and error-prone. We introduce K-step GUI Transition, a self-supervised inverse dynamics task in which VLMs learn GUI dynamics by predicting the initial action that causes a transition between two GUI states. This approach eliminates the need for natural language instructions and enables scalable dataset construction from existing GUI trajectories or automated exploration. Building on this task, we propose GUI-Shift, a reinforcement learning (RL) framework that combines rule-based optimization with data filtering to improve VLM performance. We conduct extensive experiments using multiple VLM backbones across five benchmarks, spanning GUI task automation (AndroidControl, GUI Odyssey, AndroidWorld) and GUI grounding (ScreenSpot-v2, ScreenSpot-Pro). Our results show that training on GUI-Shift generalizes well to both GUI automation and grounding tasks, yielding up to an 11.2% increase in GUI automation accuracy. This study underscores the potential of self-supervised RL to leverage unlabeled GUI trajectories and offers a scalable alternative to training with annotated samples. GUI-Shift will be open-sourced at: https://github.com/UbiquitousLearning/GUI-Shift.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper11
- CogVLM: Visual Expert for Pretrained Language ModelsWeihan Wang, Qingsong Lv, Wenmeng Yu, Wenyi Hong 等NeurIPS 2024 · 被引用 858 次
- AutoDroid: LLM-powered Task Automation in AndroidHao Wen, Yuanchun Li, Guohong Liu, Shanhui Zhao 等MobiCom 2024 · 被引用 94 次
- Inverse Dynamics Pretraining Learns Good Representations for Multitask ImitationDavid Brandfonbrener, Ofir Nachum, Joan BrunaNeurIPS 2023 · 被引用 38 次
- SeeClick: Harnessing GUI Grounding for Advanced Visual GUI AgentsKanzhi Cheng, Qiushi Sun, Yougang Chu, Fangzhi Xu 等ACL 2024 · 被引用 33 次
- ScreenSpot-Pro: GUI Grounding for Professional High-Resolution Computer UseKaixin Li, Ziyang Meng, Hongzhan Lin, Ziyang Luo 等ACM MM 2025 · 被引用 24 次
相关 Paper
- GuirlVG: Incentivize GUI Visual Grounding via Empirical Exploration on Reinforcement LearningWeitai Kang, Bin Lei, Gaowen Liu, Caiwen Ding 等ICLR 2026 · 被引用 6 次
- GUI-Eyes: Tool-Augmented Perception for Visual Grounding in GUI AgentsChen Chen, Jiawei Shao, Dakuan Lu, Haoyi Hu 等AAAI 2026 · 被引用 5 次
- Grounding Multimodal Large Language Model in GUI WorldWeixian Lei, Difei Gao, Mike Zheng ShouICLR 2025
- Test-Time Reinforcement Learning for GUI Grounding via Region ConsistencyYong Du, Yuchen Yan, Fei Tang, Zhengxi Lu 等AAAI 2026 · 被引用 10 次
- Video2GUI: Synthesizing Large-Scale Interaction Trajectories for Generalized GUI Agent PretrainingWeimin Xiong, Shuhao Gu, Bowen Ye, Zihao Yue 等ICML 2026 · 被引用 2 次
