VenusBench-Mobile: A Challenging and User-Centric Benchmark for Mobile GUI Agents with Capability Diagnostics
Yichen Gong, Zhuohan Cai, Sunhao Dai, Yuqi Zhou, Zhangxuan Gu, Changhua Meng, Shuheng Shen
Abstract
Existing online benchmarks for mobile GUI agents remain largely app-centric and taskhomogeneous, failing to reflect the diversity and instability of real-world mobile usage. To this end, we introduce VenusBench-Mobile, a challenging online benchmark for evaluating general-purpose mobile GUI agents under realistic, user-centric conditions. VenusBench-Mobile builds two core evaluation pillars: defining what to evaluate via user-intent-driven task design that reflects real mobile usage, and how to evaluate through a capability-oriented annotation scheme for fine-grained agent behavior analysis. Extensive evaluation of stateof-the-art mobile GUI agents reveals large performance gaps relative to prior benchmarks, indicating that VenusBench-Mobile poses substantially more challenging and realistic tasks and that current agents remain far from reliable real-world deployment. Diagnostic analysis further shows that failures are dominated by deficiencies in perception and memory, which are largely obscured by coarse-grained evaluations. Moreover, even the strongest agents exhibit near-zero success under environment variations, highlighting their brittleness in realistic settings. Based on these insights, we believe VenusBench-Mobile provides an important stepping stone toward robust real-world deployment of mobile GUI agents. Code and data are available at https://github.com/inclusionAI/UI-Venus/tree/VenusBench-Mobile .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8b1a5ece-214b-4807-8aa4-17a458af184aBuilds on6
- GUI-G1: Understanding R1-Zero-Like Training for Visual Grounding in GUI AgentsYuqi Zhou, Sunhao Dai, Shuai Wang, Kaiwen Zhou et al.NeurIPS 2025 · 73 citations
- AndroidLab: Training and Systematic Benchmarking of Android Autonomous AgentsYifan Xu, Xiao Liu, Xueqiao Sun, Siyi Cheng et al.ACL 2025 · 71 citations
- GUIOdyssey: A Comprehensive Dataset for Cross-App GUI Navigation on Mobile DevicesQuanfeng Lu, Wenqi Shao, Zitao Liu, Lingxiao Du et al.ICCV 2025 · 7 citations
- Aguvis: Unified Pure Vision Agents for Autonomous GUI InteractionYiheng Xu, Zekun Wang, Junli Wang, Dunjie Lu et al.ICML 2025
- MVISU-Bench: Benchmarking Mobile Agents for Real-World Tasks by Multi-App, Vague, Interactive, Single-App and Unethical InstructionsZeyu Huang, Juyuan Wang, Longfeng Chen, Boyi Xiao et al.ACM MM 2025
Related papers
- SMAN-Bench: A Cross-System Benchmark for Mobile Agents under Single- and Multi-path, Ambiguous, and Noisy TasksWeikai Xu, Zhizheng Jiang, Yuxuan Liu, Pengzhi Gao et al.ICLR 2026
- UINavBench: A Framework for Comprehensive Evaluation of Interactive Digital AgentsHarsh Agrawal, Eldon Schoop, Xinlei Pan, Anuj Mahajan et al.ICCV 2025 · 9 citations
- ProBench: Benchmarking GUI Agents with Accurate Process InformationLeyang Yang, Ziwei Wang, Xiaoxuan Tang, Sheng Zhou et al.AAAI 2026
- Mobile-Bench: An Evaluation Benchmark for LLM-based Mobile AgentsShihan Deng, Weikai Xu, Hongda Sun, Wei Liu et al.ACL 2024 · 10 citations
- FedMABench: Benchmarking Mobile GUI Agents on Decentralized Heterogeneous User DataWenhao Wang, Zijie Yu, Rui Ye, Jianqing Zhang et al.EMNLP 2025
