GUI-Bee: Align GUI Action Grounding to Novel Environments via Autonomous Exploration
Yue Fan, Handong Zhao, Ruiyi Zhang, Yu Shen, Xin Eric Wang, Gang Wu
摘要
Graphical User Interface (GUI) action grounding is a critical step in GUI automation that maps language instructions to actionable elements on GUI screens. Most recent works of GUI action grounding leverage large GUI datasets to fine-tune MLLMs. However, the fine-tuning data always covers limited GUI environments, and we find the performance of the resulting model deteriorates in novel environments. We argue that the GUI grounding models should be further aligned to the novel environments to reveal their full potential, when the inference is known to involve novel environments, i.e., environments not used during the previous fine-tuning. To realize this, we first propose GUI-Bee, an MLLM-based autonomous agent, to collect high-quality, environment-specific data through exploration and then continuously fine-tune GUI grounding models with the collected data. Our agent leverages a novel Q-value-Incentive In-Context Reinforcement Learning (Q-ICRL) method to optimize exploration efficiency and data quality. Additionally, we introduce NovelScreenSpot, a benchmark for testing how well the data can help align GUI action grounding models to novel environments and demonstrate the effectiveness of data collected by GUI-Bee in the experiments. Furthermore, we conduct an ablation study to validate the Q-ICRL method in enhancing the efficiency of GUI-Bee. Project page: https://gui-bee.github.io
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- ColorBench: Benchmarking Mobile Agents with Graph-Structured Framework for Complex Long-Horizon TasksYuanyi Song, Heyuan Huang, Qiqiang Lin, Yin Zhao 等WWW 2026 · 被引用 5 次
- Morae: Proactively Pausing UI Agents for User ChoicesYi-Hao Peng, Dingzeyu Li, Jeffrey P. Bigham, Amy PavelUIST 2025 · 被引用 4 次
- Empowering GUI Agents via Autonomous Experience Exploration and Hindsight Experience Utilization for Task PlanningTianyi Men, Zhuoran Jin, Pengfei Cao, Yubo Chen 等ACL 2026
- LongHorizonUI: A Unified Framework for Robust long-horizon Task Automation of GUI AgentBin Kang, Shaoguo Wen, Yifei Bi, Shunlong Wu 等ICLR 2026
它引用的顶会 Paper9
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- What Can Transformers Learn In-Context? A Case Study of Simple Function ClassesShivam Garg, Dimitris Tsipras, Percy Liang, Gregory ValiantNeurIPS 2022 · 被引用 883 次
- Data Distributional Properties Drive Emergent In-Context Learning in TransformersStephanie C. Y. Chan, Adam Santoro, Andrew K. Lampinen, Jane X. Wang 等NeurIPS 2022 · 被引用 407 次
- SeeClick: Harnessing GUI Grounding for Advanced Visual GUI AgentsKanzhi Cheng, Qiushi Sun, Yougang Chu, Fangzhi Xu 等ACL 2024 · 被引用 33 次
- Label Words are Anchors: An Information Flow Perspective for Understanding In-Context LearningLean Wang, Lei Li, Damai Dai, Deli Chen 等EMNLP 2023 · 被引用 26 次
相关 Paper
- Grounding Multimodal Large Language Model in GUI WorldWeixian Lei, Difei Gao, Mike Zheng ShouICLR 2025
- InfiGUI-G1: Advancing GUI Grounding with Adaptive Exploration Policy OptimizationYuhang Liu, Zeyu Liu, Shuanghe Zhu, Pengxiang Li 等AAAI 2026 · 被引用 15 次
- GuirlVG: Incentivize GUI Visual Grounding via Empirical Exploration on Reinforcement LearningWeitai Kang, Bin Lei, Gaowen Liu, Caiwen Ding 等ICLR 2026 · 被引用 6 次
- UI-Ins: Enhancing GUI Grounding with Multi-Perspective Instruction as ReasoningLiangyu Chen, Hanzhang Zhou, chenglin Cai, Jianan Zhang 等ICLR 2026 · 被引用 19 次
- Continual GUI AgentsZiwei Liu, Borui Kang, Hangjie Yuan, Zixiang Zhao 等ICML 2026 · 被引用 6 次
