MobileWorld: Benchmarking Autonomous Mobile Agents in Agent-User Interactive and MCP-Augmented Environments
Quyu Kong, Xu Zhang, Zhenyu Yang, Nolan Gao, Chen Liu, Panrong Tong, Chenglin Cai, Hanzhang Zhou, Jianan Zhang, Liangyu Chen, Zhidan Liu, Steven Hoi, Yue Wang
摘要
Among existing online mobile-use benchmarks, AndroidWorld has emerged as the dominant benchmark due to its reproducible environment and deterministic evaluation; however, recent agents achieving over 90% success rates indicate its saturation and motivate the need for a more challenging benchmark. In addition, its environment lacks key application categories, such as e-commerce and enterprise communication, and does not reflect realistic mobile-use scenarios characterized by vague user instructions and hybrid tool usage. To bridge these gaps, we introduce MobileWorld, a substantially more challenging benchmark designed to reflect real-world usage through 201 tasks across 20 applications. MobileWorld derives its difficulty from an emphasis on long-horizon, cross-application workflows, requiring nearly twice as many completion steps on average (27.8 vs. 14.3) and featuring a significantly higher proportion of multi-app tasks (62.2% vs. 9.5%) than AndroidWorld. To overcome the limitations of existing environments, MobileWorld achieves a balance between production-grade utility and reproducible evaluation by utilizing open-source alternatives to industry standards (e.g., Mattermost for Slack). This approach enables a fully observable and controlled environment through source code modification and direct backend database access for precise verification. Furthermore, MobileWorld extends beyond standard GUI manipulation by introducing novel task categories, including agent-user interaction and Model Context Protocol (MCP)-augmented tasks, providing a robust framework for evaluating agents in user-aware, hybrid-tool scenarios. To facilitate evaluation, we develop a planner-executor agentic framework with extended action spaces to support user interactions and MCP calls. Our results reveal a sharp performance drop compared to AndroidWorld, with the best agentic framework and end-to-end model achieving 51.7% and 20.9% success rates, respectively, highlighting ample headroom for future research. Our analysis further shows that current models struggle significantly with user interaction and MCP calls. By identifying these core research gaps, we offer a strategic roadmap toward next-generation mobile intelligence.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- InfiniteWeb: Scalable Web Environment Synthesis for GUI Agent TrainingZiyun Zhang, Zezhou Wang, Xiaoyi Zhang, Zongyu Guo 等ACL 2026 · 被引用 11 次
- MaDS: Long-Horizon GUI Automation via Synergizing Dual-Layer Memory and Multi-Round DebatePengchen Chen, Shi Chen, Qiming Ye, Xinli Chen 等ACL 2026
- UIAnchor: Anchoring UI Perception and Action Execution for Reliable Service-Composed Mobile Task Automation with GUI AgentsWentao Zhou, Sicong Liu, Zimu Zhou, Yimeng Duan 等UbiComp 2026
它引用的顶会 Paper8
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
- AndroidLab: Training and Systematic Benchmarking of Android Autonomous AgentsYifan Xu, Xiao Liu, Xueqiao Sun, Siyi Cheng 等ACL 2025 · 被引用 71 次
- OSWorld-MCP: Benchmarking MCP Tool Invocation In Computer-Use AgentsHongrui Jia, Jitong Liao, Xi Zhang, Haiyang Xu 等ICLR 2026 · 被引用 24 次
- UI-Ins: Enhancing GUI Grounding with Multi-Perspective Instruction as ReasoningLiangyu Chen, Hanzhang Zhou, chenglin Cai, Jianan Zhang 等ICLR 2026 · 被引用 19 次
- Mobile-Bench: An Evaluation Benchmark for LLM-based Mobile AgentsShihan Deng, Weikai Xu, Hongda Sun, Wei Liu 等ACL 2024 · 被引用 10 次
相关 Paper
- AndroidWorld: A Dynamic Benchmarking Environment for Autonomous AgentsChristopher Rawles, Sarah Clinckemaillie, Yifan Chang, Jonathan Waltz 等ICLR 2025
- MCP-AgentBench: Evaluating Real-World Language Agent Performance with MCP-Mediated ToolsZikang Guo, Benfeng Xu, Chiwei Zhu, Wentao Hong 等AAAI 2026 · 被引用 19 次
- MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP ServersZhenting Wang, Qi Chang, Hemani Patel, Shashank Biju 等ICLR 2026 · 被引用 109 次
- macOSWorld: A Multilingual Interactive Benchmark for GUI AgentsPei Yang, Hai Ci, Mike Zheng ShouNeurIPS 2025 · 被引用 34 次
- ComplexMCP: Evaluation of LLM Agents in Dynamic, Interdependent, and Large-Scale Tool SandboxYuanyang Li, Xue Yang, Longyue Wang, Weihua Luo 等ICML 2026 · 被引用 1 次
