Don't Click That: Teaching Web Agents to Resist Deceptive Interfaces
Yilin Zhang, Yingkai Hua, Chunyu Wei, Xin Wang, Yueguo Chen
Abstract
Vision-language model (VLM) based web agents demonstrate impressive autonomous GUI interaction but remain vulnerable to deceptive interface elements. Existing approaches either detect deception without task integration or document attacks without proposing defenses. We formalize deception-aware web agent defense and propose DUDE (Deceptive UI Detector & Evaluator), a two-stage framework combining hybrid-reward learning with asymmetric penalties and experience summarization to distill failure patterns into transferable guidance. We introduce RUC (Real UI Clickboxes), a benchmark of 1,407 scenarios spanning four domains and deception categories. Experiments show DUDE reduces deception susceptibility by 53.8% while maintaining task performance, establishing an effective foundation for robust web agent deployment. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bdb0e105-4891-4ba9-9dbf-b8dc3c37cb81Builds on16
- WebArena: A Realistic Web Environment for Building Autonomous AgentsShuyan Zhou, Frank F. Xu, Hao Zhu, Xuhui Zhou et al.ICLR 2024 · 1,197 citations
- Attacking Vision-Language Computer Agents via Pop-upsYanzhe Zhang, Tao Yu, Diyi YangACL 2025 · 99 citations
- Fight Back Against Jailbreaking via Prompt Adversarial TuningYichuan Mo, Yuji Wang, Zeming Wei, Yisen WangNeurIPS 2024 · 90 citations
- RedTeamCUA: Realistic Adversarial Testing of Computer-Use Agents in Hybrid Web-OS EnvironmentsZeyi Liao, Jaylen Jones, Linxi Jiang, Yuting Ning et al.ICLR 2026 · 46 citations
- Unveiling the Tricks: Automated Detection of Dark Patterns in Mobile ApplicationsJieshan Chen, Jiamou Sun, Sidong Feng, Zhenchang Xing et al.UIST 2023 · 42 citations
Related papers
- Benchmarking Web Agent Safety under E-commerce Deceptive InterfacesZijing Shi, Meng Fang, Ling ChenACL 2026
- VPI-Bench: Visual Prompt Injection Attacks for Computer-Use AgentsTri Cao, Bennett Lim, Yue Liu, Yuan Sui et al.ICLR 2026 · 45 citations
- Automatically Detecting Online Deceptive PatternsAsmit Nayak, Yash Wani, Shirley Zhang, Rishabh Khandelwal et al.CCS 2025
- How Dark Patterns Manipulate Web AgentsPhil Cuvin, Hao Zhu, Diyi YangICLR 2026 · 9 citations
- 50 Shades of Deceptive Patterns: A Unified Taxonomy, Multimodal Detection, and Security ImplicationsZewei Shi, Ruoxi Sun, Jieshan Chen, Jiamou Sun et al.WWW 2025 · 12 citations
