MMBench-GUI: A Unified Hierarchical Evaluation Framework for Multi-Platform GUI Agents
Xuehui Wang, Zhenyu Wu, JingJing Xie, Zichen Ding, Bowen Yang, Zehao Li, Zhaoyang Liu, Qingyun Li, Xuan Dong, Zhe Chen, Weiyun Wang, Xiangyu Zhao
摘要
We introduce MMBench-GUI, a hierarchical benchmark for evaluating GUI automation agents across Windows, ma-cOS, Linux, iOS, Android, and Web. The benchmark spans four levels: Content Understanding, Element Grounding, Task Automation, and Task Collaboration, covering essential skills for GUI agents. To assess both effectiveness and efficiency, we further propose the Efficiency-Quality-Aware (EQA) metric, which measures task success alongside action redundancy. Extensive evaluations reveal that precise visual grounding is the critical determinant of performance, underscoring the advantages of modular designs with specialized grounding modules. Moreover, all agents suffer from substantial inefficiencies, frequently completing tasks with excessive steps despite eventual success. Performance also degrades on complex or cross-application tasks, exposing weaknesses in memory, planning, and adaptive reasoning. By providing broad coverage, standardized protocols, and novel metrics, MMBench-GUI establishes the first comprehensive foundation for advancing GUI agent research. Our benchmark code, evaluation data, and running environment are publicly available at https://github.com/open-compass/MMBench-GUI.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper15
- WebArena: A Realistic Web Environment for Building Autonomous AgentsShuyan Zhou, Frank F. Xu, Hao Zhu, Xuhui Zhou 等ICLR 2024 · 被引用 1,197 次
- GPT-4V(ision) is a Generalist Web Agent, if GroundedBoyuan Zheng, Boyu Gou, Jihyung Kil, Huan Sun 等ICML 2024 · 被引用 496 次
- AndroidLab: Training and Systematic Benchmarking of Android Autonomous AgentsYifan Xu, Xiao Liu, Xueqiao Sun, Siyi Cheng 等ACL 2025 · 被引用 71 次
- ScaleCUA: Scaling Open-Source Computer Use Agents with Cross-Platform DataZhaoyang Liu, Jingjing Xie, Zichen Ding, Zehao Li 等ICLR 2026 · 被引用 54 次
- ScienceBoard: Evaluating Multimodal Autonomous Agents in Realistic Scientific WorkflowsQiushi Sun, Zhoumianze Liu, Chang Ma, Zichen Ding 等ICLR 2026 · 被引用 45 次
相关 Paper
- UINavBench: A Framework for Comprehensive Evaluation of Interactive Digital AgentsHarsh Agrawal, Eldon Schoop, Xinlei Pan, Anuj Mahajan 等ICCV 2025 · 被引用 9 次
- DAC-Bench: A Decision-Aware Benchmark for Compositional Mobile GUI TasksYuqing Zhang, Honghui Sheng, Xueyu Hu, Shengyu Zhang 等ACL 2026
- GUI-CEval: A Hierarchical and Comprehensive Chinese Benchmark for Mobile GUI AgentsYang Li, Yuchen Liu, Haoyu Lu, Zhiqiang Xia 等CVPR 2026 · 被引用 3 次
- GUIDE: A Benchmark for Understanding and Assisting Users in Open-Ended GUI TasksSaelyne Yang, Jaesang Yu, Yi-Hao Peng, Kevin Qinghong Lin 等CVPR 2026 · 被引用 5 次
- Grounding Multimodal Large Language Model in GUI WorldWeixian Lei, Difei Gao, Mike Zheng ShouICLR 2025
