Don't Click That: Teaching Web Agents to Resist Deceptive Interfaces
Yilin Zhang, Yingkai Hua, Chunyu Wei, Xin Wang, Yueguo Chen
摘要
Vision-language model (VLM) based web agents demonstrate impressive autonomous GUI interaction but remain vulnerable to deceptive interface elements. Existing approaches either detect deception without task integration or document attacks without proposing defenses. We formalize deception-aware web agent defense and propose DUDE (Deceptive UI Detector & Evaluator), a two-stage framework combining hybrid-reward learning with asymmetric penalties and experience summarization to distill failure patterns into transferable guidance. We introduce RUC (Real UI Clickboxes), a benchmark of 1,407 scenarios spanning four domains and deception categories. Experiments show DUDE reduces deception susceptibility by 53.8% while maintaining task performance, establishing an effective foundation for robust web agent deployment. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper16
- WebArena: A Realistic Web Environment for Building Autonomous AgentsShuyan Zhou, Frank F. Xu, Hao Zhu, Xuhui Zhou 等ICLR 2024 · 被引用 1,197 次
- Attacking Vision-Language Computer Agents via Pop-upsYanzhe Zhang, Tao Yu, Diyi YangACL 2025 · 被引用 99 次
- Fight Back Against Jailbreaking via Prompt Adversarial TuningYichuan Mo, Yuji Wang, Zeming Wei, Yisen WangNeurIPS 2024 · 被引用 90 次
- RedTeamCUA: Realistic Adversarial Testing of Computer-Use Agents in Hybrid Web-OS EnvironmentsZeyi Liao, Jaylen Jones, Linxi Jiang, Yuting Ning 等ICLR 2026 · 被引用 46 次
- Unveiling the Tricks: Automated Detection of Dark Patterns in Mobile ApplicationsJieshan Chen, Jiamou Sun, Sidong Feng, Zhenchang Xing 等UIST 2023 · 被引用 42 次
相关 Paper
- Benchmarking Web Agent Safety under E-commerce Deceptive InterfacesZijing Shi, Meng Fang, Ling ChenACL 2026
- VPI-Bench: Visual Prompt Injection Attacks for Computer-Use AgentsTri Cao, Bennett Lim, Yue Liu, Yuan Sui 等ICLR 2026 · 被引用 45 次
- Automatically Detecting Online Deceptive PatternsAsmit Nayak, Yash Wani, Shirley Zhang, Rishabh Khandelwal 等CCS 2025
- How Dark Patterns Manipulate Web AgentsPhil Cuvin, Hao Zhu, Diyi YangICLR 2026 · 被引用 9 次
- 50 Shades of Deceptive Patterns: A Unified Taxonomy, Multimodal Detection, and Security ImplicationsZewei Shi, Ruoxi Sun, Jieshan Chen, Jiamou Sun 等WWW 2025 · 被引用 12 次
