UI-Vision: A Desktop-centric GUI Benchmark for Visual Perception and Interaction
Shravan Nayak, Xiangru Jian, Kevin Qinghong Lin, Juan A. Rodríguez, Montek Kalsi, Nicolas Chapados, M. Tamer Özsu, Aishwarya Agrawal, David Vázquez, Christopher Pal, Perouz Taslakian, Spandana Gella, Sai Rajeswar
摘要
Autonomous agents that navigate Graphical User Interfaces (GUIs) to automate tasks like document editing and file management can greatly enhance computer workflows. While existing research focuses on online settings, desktop environments, critical for many professional and everyday tasks, remain underexplored due to data collection challenges and licensing issues. We introduce UI-Vision, the first comprehensive, licensepermissive benchmark for offline, fine-grained evaluation of computer use agents in real-world desktop environments. Unlike online benchmarks, UI-Vision provides: (i) dense, high-quality annotations of human demonstrations, including bounding boxes, UI labels, and action trajectories (clicks, drags, and keyboard inputs) across 83 software applications, and (ii) three fine-tocoarse grained tasks-Element Grounding, Layout Grounding, and Action Prediction-with well-defined metrics to rigorously evaluate agents' performance in desktop environments. Our evaluation reveals critical limitations in state-of-the-art models like UI-TARS-72B, including issues with understanding professional software, spatial reasoning, and complex actions like drag-and-drop. These findings highlight the challenges in developing fully autonomous computer-use agents. With UI-Vision, we aim to advance the development of more capable agents for real-world desktop tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper16
- OpenCUA: Open Foundations for Computer-Use AgentsXinyuan Wang, Bowen Wang, Dunjie Lu, Junlin Yang 等NeurIPS 2025 · 被引用 151 次
- GTA1: GUI Test-time Scaling AgentYan Yang, Dongxu Li, Yutong Dai, Yuhao Yang 等ICLR 2026 · 被引用 109 次
- Rendering-Aware Reinforcement Learning for Vector Graphics GenerationJuan A. Rodríguez, Haotian Zhang, Abhay Puri, Rishav Pramanik 等NeurIPS 2025 · 被引用 42 次
- MVP: Multiple View Prediction Improves GUI GroundingYunzhu Zhang, Zeyu Pan, Zhengwen Zeng, Shuheng Shen 等CVPR 2026 · 被引用 10 次
- Learning GUI Grounding with Spatial Reasoning from Visual FeedbackYu Zhao, Wei-Ning Chen, Huseyin Inan, Samuel Kessler 等ICML 2026 · 被引用 9 次
它引用的顶会 Paper16
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Toolformer: Language Models Can Teach Themselves to Use ToolsTimo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu 等NeurIPS 2023 · 被引用 5,989 次
- WebShop: Towards Scalable Real-World Web Interaction with Grounded Language AgentsShunyu Yao, Howard Chen, John Yang, Karthik NarasimhanNeurIPS 2022 · 被引用 1,477 次
- GPT-4V(ision) is a Generalist Web Agent, if GroundedBoyuan Zheng, Boyu Gou, Jihyung Kil, Huan Sun 等ICML 2024 · 被引用 496 次
相关 Paper
- Aguvis: Unified Pure Vision Agents for Autonomous GUI InteractionYiheng Xu, Zekun Wang, Junli Wang, Dunjie Lu 等ICML 2025
- ShowUI-π: Flow-based Generative Models as GUI Dexterous HandsSiyuan Hu, Kevin Qinghong Lin, Mike Zheng ShouCVPR 2026 · 被引用 4 次
- GUIDE: A Benchmark for Understanding and Assisting Users in Open-Ended GUI TasksSaelyne Yang, Jaesang Yu, Yi-Hao Peng, Kevin Qinghong Lin 等CVPR 2026 · 被引用 5 次
- MMBench-GUI: A Unified Hierarchical Evaluation Framework for Multi-Platform GUI AgentsXuehui Wang, Zhenyu Wu, JingJing Xie, Zichen Ding 等CVPR 2026
- VenusBench-Mobile: A Challenging and User-Centric Benchmark for Mobile GUI Agents with Capability DiagnosticsYichen Gong, Zhuohan Cai, Sunhao Dai, Yuqi Zhou 等ICML 2026 · 被引用 1 次
