GTA1: GUI Test-time Scaling Agent
Yan Yang, Dongxu Li, Yutong Dai, Yuhao Yang, Ziyang Luo, Zirui Zhao, Zhiyuan Hu, Junzhe Huang, Amrita Saha, Zeyuan Chen, Ran Xu, Liyuan Pan
摘要
Graphical user interface (GUI) agents autonomously complete tasks across platforms (, Linux) by sequentially decomposing user instructions into action proposals that iteratively interact with visual elements in the evolving environment. However, two main challenges arise: i) planning (, the action proposal sequence) under expansive action space, where selecting an appropriate plan is non-trivial, as many valid ones may exist; ii) accurately grounding actions in complex and high-resolution interfaces, , precisely interacting with visual targets. This paper investigates the aforementioned challenges with our GUI Test-time Scaling Agent, namely GTA1. First, we conduct test-time scaling to select the most appropriate action proposal: at each step, multiple candidate proposals are sampled and evaluated and selected by a judge model. It trades off computation for better decision quality by concurrent sampling. Second, we propose a model that improves grounding of the selected action proposals to its corresponding visual elements. Our key insight is that reinforcement learning (RL) facilitates grounding through inherent objective alignments, rewarding successful clicks on interface elements. Experimentally, GTA1 achieves state-of-the-art performance on both grounding and agent task execution benchmarks. The code and models are released here.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper23
- ScaleCUA: Scaling Open-Source Computer Use Agents with Cross-Platform DataZhaoyang Liu, Jingjing Xie, Zichen Ding, Zehao Li 等ICLR 2026 · 被引用 54 次
- CoAct-1: Computer-using Multi-agent System with Coding ActionsLinxin Song, Yutong Dai, Viraj Prabhu, Jieyu Zhang 等ICLR 2026 · 被引用 32 次
- UI-Ins: Enhancing GUI Grounding with Multi-Perspective Instruction as ReasoningLiangyu Chen, Hanzhang Zhou, chenglin Cai, Jianan Zhang 等ICLR 2026 · 被引用 19 次
- RealWebAssist: A Benchmark for Long-Horizon Web Assistance with Real-World UsersSuyu Ye, Haojun Shi, Darren Shih, Hyokun Yun 等AAAI 2026 · 被引用 17 次
- Grounding Computer Use Agents on Human DemonstrationsAarash Feizi, Shravan Nayak, Xiangru Jian, Kevin Qinghong Lin 等ICLR 2026 · 被引用 16 次
它引用的顶会 Paper15
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
- Is ChatGPT a General-Purpose Natural Language Processing Task Solver?Chengwei Qin, Aston Zhang, Zhuosheng Zhang, Jiaao Chen 等EMNLP 2023 · 被引用 449 次
- OpenCUA: Open Foundations for Computer-Use AgentsXinyuan Wang, Bowen Wang, Dunjie Lu, Junlin Yang 等NeurIPS 2025 · 被引用 151 次
- GUI-G1: Understanding R1-Zero-Like Training for Visual Grounding in GUI AgentsYuqi Zhou, Sunhao Dai, Shuai Wang, Kaiwen Zhou 等NeurIPS 2025 · 被引用 73 次
- Widget Captioning: Generating Natural Language Description for Mobile User Interface ElementsYang Li, Gang Li, Luheng He, Jingjie Zheng 等EMNLP 2020 · 被引用 46 次
相关 Paper
- Test-Time Reinforcement Learning for GUI Grounding via Region ConsistencyYong Du, Yuchen Yan, Fei Tang, Zhengxi Lu 等AAAI 2026 · 被引用 10 次
- Learning GUI Grounding with Spatial Reasoning from Visual FeedbackYu Zhao, Wei-Ning Chen, Huseyin Inan, Samuel Kessler 等ICML 2026 · 被引用 9 次
- Visual Test-Time Scaling for GUI Agent GroundingTiange Luo, Lajanugen Logeswaran, Justin Johnson, Honglak LeeICCV 2025 · 被引用 3 次
- SeeClick: Harnessing GUI Grounding for Advanced Visual GUI AgentsKanzhi Cheng, Qiushi Sun, Yougang Chu, Fangzhi Xu 等ACL 2024 · 被引用 33 次
- MMBench-GUI: A Unified Hierarchical Evaluation Framework for Multi-Platform GUI AgentsXuehui Wang, Zhenyu Wu, JingJing Xie, Zichen Ding 等CVPR 2026
