Grounding Computer Use Agents on Human Demonstrations
Aarash Feizi, Shravan Nayak, Xiangru Jian, Kevin Qinghong Lin, Kaixin Li, Rabiul Awal, Xing Han Lù, Johan S. Obando Ceron, Juan A. Rodríguez, Nicolas Chapados, David Vázquez, Adriana Romero-Soriano
摘要
Building reliable computer-use agents requires grounding: accurately connecting natural language instructions to the correct on-screen elements. While large datasets exist for web and mobile interactions, high-quality resources for desktop environments are limited. To address this gap, we introduce GROUNDCUA, a large-scale desktop grounding dataset built from expert human demonstrations. It covers 87 applications across 12 categories and includes 56K screenshots, with every on-screen element carefully annotated for a total of over 3.56M humanverified annotations. From these demonstrations, we generate diverse instructions that capture a wide range of real-world tasks, providing high-quality data for model training. Using GROUNDCUA, we develop the GROUNDNEXT family of models that map instructions to their target UI elements. At both 3B and 7B scales, GROUNDNEXT achieves state-of-the-art results across five benchmarks using supervised fine-tuning, while requiring less than one-tenth the training data of prior work. Reinforcement learning post-training further improves performance, and when evaluated in an agentic setting on the OSWorld benchmark using o3 as planner, GROUNDNEXT attains comparable or superior results to models trained with substantially more data,. These results demonstrate the critical role of highquality, expert-driven datasets in advancing general-purpose computer-use agents. * Equal contribution. † Equal supervision.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Rendering-Aware Reinforcement Learning for Vector Graphics GenerationJuan A. Rodríguez, Haotian Zhang, Abhay Puri, Rishav Pramanik 等NeurIPS 2025 · 被引用 42 次
- GUIDE: A Benchmark for Understanding and Assisting Users in Open-Ended GUI TasksSaelyne Yang, Jaesang Yu, Yi-Hao Peng, Kevin Qinghong Lin 等CVPR 2026 · 被引用 5 次
- ScreenParse: Moving Beyond Sparse Grounding with Complete Screen Parsing SupervisionA. Said Gurbuz, Sunghwan Hong, Ahmed Nassar, Marc Pollefeys 等ICML 2026 · 被引用 3 次
- TreeCUA: Efficiently Scaling GUI Automation with Tree-Structured Verifiable EvolutionDeyang Jiang, Jing Huang, Xuanle Zhao, Lei Chen 等ICML 2026
它引用的顶会 Paper4
- GTA1: GUI Test-time Scaling AgentYan Yang, Dongxu Li, Yutong Dai, Yuhao Yang 等ICLR 2026 · 被引用 109 次
- GUI-G1: Understanding R1-Zero-Like Training for Visual Grounding in GUI AgentsYuqi Zhou, Sunhao Dai, Shuai Wang, Kaiwen Zhou 等NeurIPS 2025 · 被引用 73 次
- SeeClick: Harnessing GUI Grounding for Advanced Visual GUI AgentsKanzhi Cheng, Qiushi Sun, Yougang Chu, Fangzhi Xu 等ACL 2024 · 被引用 33 次
- Back to Basics: Revisiting REINFORCE-Style Optimization for Learning from Human Feedback in LLMsArash Ahmadian, Chris Cremer, Matthias Gallé, Marzieh Fadaee 等ACL 2024 · 被引用 20 次
相关 Paper
- Anchor: Branch-Point Data Generation for GUI AgentsJinbiao Wei, Yilun Zhao, Kangqi Ni, Arman CohanACL 2026 · 被引用 3 次
- ScaleCUA: Scaling Open-Source Computer Use Agents with Cross-Platform DataZhaoyang Liu, Jingjing Xie, Zichen Ding, Zehao Li 等ICLR 2026 · 被引用 54 次
- OpenCUA: Open Foundations for Computer-Use AgentsXinyuan Wang, Bowen Wang, Dunjie Lu, Junlin Yang 等NeurIPS 2025 · 被引用 151 次
- Watch and Learn: Learning to Use Computers from Online VideosChan Hee Song, Yiwen Song, Palash Goyal, Yu Su 等CVPR 2026 · 被引用 8 次
- VideoAgentTrek: Computer-Use Pretraining from Unlabeled VideosDunjie Lu, Yiheng Xu, Junli Wang, Haoyuan Wu 等ICLR 2026 · 被引用 20 次
