macOSWorld: A Multilingual Interactive Benchmark for GUI Agents
Pei Yang, Hai Ci, Mike Zheng Shou
摘要
Graphical User Interface (GUI) agents show promising capabilities for automating computer-use tasks and facilitating accessibility, but existing interactive benchmarks are mostly English-only, covering web-use or Windows, Linux, and Android environments, but not macOS. macOS is a major OS with distinctive GUI patterns and exclusive applications. To bridge the gaps, we present macOSWorld, the first comprehensive benchmark for evaluating GUI agents on macOS. macOSWorld features 202 multilingual interactive tasks across 30 applications (28 macOS-exclusive), with task instructions and OS interfaces offered in 5 languages (English, Chinese, Arabic, Japanese, and Russian). As GUI agents are shown to be vulnerable to deception attacks, macOSWorld also includes a dedicated safety benchmarking subset. Our evaluation on six GUI agents reveals a dramatic gap: proprietary computer-use agents lead at above 30% success rate, while open-source lightweight research models lag at below 5%, highlighting the need for macOS domain adaptation. Multilingual benchmarks also expose common weaknesses, especially in Arabic, with a 28.8% average degradation compared to English. Results from safety benchmarking also highlight that deception attacks are more general and demand immediate attention. Project page: https://macos-world.github.io.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- SCUBA: Salesforce Computer Use BenchmarkYutong Dai, Krithika Ramakrishnan, Jing Gu, Matthew Fernandez 等ICLR 2026 · 被引用 8 次
- MPR-GUI: Benchmarking and Enhancing Multilingual Perception and Reasoning in GUI AgentsRuihan Chen, Qiming Li, Xiaocheng Feng, Weihong Zhong 等ACL 2026 · 被引用 3 次
它引用的顶会 Paper28
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- WebArena: A Realistic Web Environment for Building Autonomous AgentsShuyan Zhou, Frank F. Xu, Hao Zhu, Xuhui Zhou 等ICLR 2024 · 被引用 1,197 次
- GAIA: a benchmark for General AI AssistantsGrégoire Mialon, Clémentine Fourrier, Thomas Wolf, Yann LeCun 等ICLR 2024 · 被引用 716 次
- AgentPoison: Red-teaming LLM Agents via Poisoning Memory or Knowledge BasesZhaorun Chen, Zhen Xiang, Chaowei Xiao, Dawn Song 等NeurIPS 2024 · 被引用 539 次
- GPT-4V(ision) is a Generalist Web Agent, if GroundedBoyuan Zheng, Boyu Gou, Jihyung Kil, Huan Sun 等ICML 2024 · 被引用 496 次
相关 Paper
- AndroidWorld: A Dynamic Benchmarking Environment for Autonomous AgentsChristopher Rawles, Sarah Clinckemaillie, Yifan Chang, Jonathan Waltz 等ICLR 2025
- HackWorld: Evaluating Computer-Use Agents on Exploiting Web Application VulnerabilitiesXiaoxue Ren, Penghao Jiang, Kaixin Li, Zhiyong Huang 等ICLR 2026 · 被引用 3 次
- MobileWorld: Benchmarking Autonomous Mobile Agents in Agent-User Interactive and MCP-Augmented EnvironmentsQuyu Kong, Xu Zhang, Zhenyu Yang, Nolan Gao 等ACL 2026 · 被引用 38 次
- RiOSWorld: Benchmarking the Risk of Multimodal Computer-Use AgentsJingyi Yang, Shuai Shao, Dongrui Liu, Jing ShaoNeurIPS 2025 · 被引用 33 次
- VPI-Bench: Visual Prompt Injection Attacks for Computer-Use AgentsTri Cao, Bennett Lim, Yue Liu, Yuan Sui 等ICLR 2026 · 被引用 45 次
