SMAN-Bench: A Cross-System Benchmark for Mobile Agents under Single- and Multi-path, Ambiguous, and Noisy Tasks
Weikai Xu, Zhizheng Jiang, Yuxuan Liu, Pengzhi Gao, Wei Liu, Jian Luan, Yunxin Liu, Yuanchun Li, Bin Wang, Bo An
Abstract
VLM-based mobile agents are increasingly popular due to their capabilities to interact with smartphone GUIs and XML-structured texts and complete daily tasks. However, existing online benchmarks fail to provide stable fine-grained reward signals under dynamic environmental changes and neglect the influence of noisy components and interactive instructions. Offline benchmarks evaluate the agents through single-path trajectories, which stand in contrast to the inherently multisolution characteristics of GUI tasks. To address these limitations, we introduce SMAN-Bench, a benchmark designed to evaluate agents under Single-path, Multipath, Ambiguous, and Noisy task settings. We employ a slot-based instruction generation method to match templates with GUI trajectories from an existing, graphstructured, unlabeled mobile corpus. SMAN-Bench includes a common task split, with offline multi-path evaluation to assess the agent's ability to obtain step rewards during task execution. It contains a noisy split based on pop-ups and ad apps, and a contaminated split to simulate a realistic noisy environment. Furthermore, an ambiguous instruction split with preset Q&A interactions is released to evaluate the agent's proactive interaction capabilities. Our evaluation encompasses mobile agent frameworks such as AppAgent-v1, Mobile-Agent-v2, and Mobile-Agent-E, and includes both open-source and closed-source mobile foundation models, as well as several multimodal thinking models. Our code and datasets are available at https://github.com/gezelligheid0314/SMAN-Bench . Our evaluation of general-purpose multimodal models is conducted on three agent frameworks:
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 256cbabc-cb55-4484-ba05-b77d7173cd87Builds on27
- GPT-4V(ision) is a Generalist Web Agent, if GroundedBoyuan Zheng, Boyu Gou, Jihyung Kil, Huan Sun et al.ICML 2024 · 496 citations
- Mobile-Agent-v2: Mobile Device Operation Assistant with Effective Navigation via Multi-Agent CollaborationJunyang Wang, Haiyang Xu, Haitao Jia, Xi Zhang et al.NeurIPS 2024 · 245 citations
- DigiRL: Training In-The-Wild Device-Control Agents with Autonomous Reinforcement LearningHao Bai, Yifei Zhou, Jiayi Pan, Mert Cemri et al.NeurIPS 2024 · 239 citations
- Enabling Conversational Interaction with Mobile UI using Large Language ModelsBryan Wang, Gang Li, Yang LiCHI 2023 · 149 citations
- UI-R1: Enhancing Efficient Action Prediction of GUI Agents by Reinforcement LearningZhengxi Lu, Yuxiang Chai, Yaxuan Guo, Xi Yin et al.AAAI 2026 · 103 citations
Related papers
- Spa-Bench: a comprehensive Benchmark for Smartphone Agent EvaluationJingxuan Chen, Derek Yuen, Bin Xie, Yuhao Yang et al.ICLR 2025
- Mobile-Bench: An Evaluation Benchmark for LLM-based Mobile AgentsShihan Deng, Weikai Xu, Hongda Sun, Wei Liu et al.ACL 2024 · 10 citations
- ColorBench: Benchmarking Mobile Agents with Graph-Structured Framework for Complex Long-Horizon TasksYuanyi Song, Heyuan Huang, Qiqiang Lin, Yin Zhao et al.WWW 2026 · 5 citations
- VenusBench-Mobile: A Challenging and User-Centric Benchmark for Mobile GUI Agents with Capability DiagnosticsYichen Gong, Zhuohan Cai, Sunhao Dai, Yuqi Zhou et al.ICML 2026 · 1 citation
- AndroidLab: Training and Systematic Benchmarking of Android Autonomous AgentsYifan Xu, Xiao Liu, Xueqiao Sun, Siyi Cheng et al.ACL 2025 · 71 citations
