Mind the Third Eye! Benchmarking Privacy Awareness in MLLM-powered Smartphone Agents
Zhixin Lin, Jungang Li, Shidong Pan, Yibo Shi, Yue Yao, Dongliang Xu
摘要
Smartphones bring significant convenience to users but also enable devices to extensively record various types of personal information. Existing smartphone agents powered by Multimodal Large Language Models (MLLMs) have achieved remarkable performance in automating different tasks. However, as the cost, these agents are granted substantial access to sensitive users' personal information during this operation. To gain a thorough understanding of the privacy awareness of these agents, we present the first large-scale benchmark encompassing 7,138 scenarios to the best of our knowledge. In addition, for privacy context in scenarios, we annotate its type (e.g., Account Credentials), sensitivity level, and location. We then carefully benchmark seven available mainstream smartphone agents. Our results demonstrate that almost all benchmarked agents show unsatisfying privacy awareness (RA), with performance remaining below 60% even with explicit hints. Overall, closed-source agents show better privacy ability than open-source ones, and Gemini 2.0-flash achieves the best, achieving an RA of 67%. We also find that the agents’ privacy detection capability is highly related to scenario sensitivity level, i.e., the scenario with a higher sensitivity level is typically more identifiable. We hope the findings enlighten the research community to rethink the unbalanced utility-privacy tradeoff about smartphone agents.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper8
- AndroidLab: Training and Systematic Benchmarking of Android Autonomous AgentsYifan Xu, Xiao Liu, Xueqiao Sun, Siyi Cheng 等ACL 2025 · 被引用 71 次
- Analyzing User Perspectives on Mobile App Privacy at ScalePreksha Nema, Pauline Anthonysamy, Nina Taft, Sai Teja PeddintiICSE 2022 · 被引用 50 次
- A NEW HOPE: Contextual Privacy Policies for Mobile Applications and An Approach Toward Automated GenerationShidong Pan, Zhen Tao, Thong Hoang, Dawen Zhang 等USENIX Security 2024 · 被引用 24 次
- Mobile-Bench: An Evaluation Benchmark for LLM-based Mobile AgentsShihan Deng, Weikai Xu, Hongda Sun, Wei Liu 等ACL 2024 · 被引用 10 次
- GUIOdyssey: A Comprehensive Dataset for Cross-App GUI Navigation on Mobile DevicesQuanfeng Lu, Wenqi Shao, Zitao Liu, Lingxiao Du 等ICCV 2025 · 被引用 7 次
相关 Paper
- Measuring Physical-World Privacy Awareness of Large Language Models: An Evaluation BenchmarkXinjie Shen, Mufei Li, Pan LiICLR 2026 · 被引用 3 次
- Spa-Bench: a comprehensive Benchmark for Smartphone Agent EvaluationJingxuan Chen, Derek Yuen, Bin Xie, Yuhao Yang 等ICLR 2025
- FingerTip 20K: A Benchmark for Proactive and Personalized Mobile LLM AgentsQinglong Yang, Haoming Li, Haotian Zhao, Xiaokai Yan 等ICLR 2026 · 被引用 21 次
- VoxPrivacy: A Benchmark for Evaluating Interactional Privacy of Speech Language ModelsYuxiang Wang, HongYu Liu, Dekun Chen, Xueyao Zhang 等ICLR 2026 · 被引用 5 次
- AgencyBench: Benchmarking the Frontiers of Autonomous Agents in 1M-Token Real-World ContextsKeyu Li, Junhao Shi, Yang Xiao, Mohan Jiang 等ACL 2026 · 被引用 14 次
