AgentBreaker: Evaluating Context-Aware Indirect Prompt Injection Risks in Modern Web Agents
Yongbi Son, Changoo Lee, Dongwon Shin, Byoungyoung Lee, Sanghyun Hong, Sooel Son
摘要
Recent advances in large language models (LLMs) have enabled autonomous web agents to perform complex user tasks by leveraging their adaptive decision-making capabilities. Despite their growing use in crawling the Web, their security implications under indirect prompt injection (IPI) attacks remain largely understudied. Prior studies have compiled static benchmarks or proposed dynamic frameworks that generate adversarial phrases aimed at deceiving a single LLM within a target agent. However, by ignoring the agent’s operating context, these approaches yield suboptimal IPI attacks against modern web agents leveraging multiple, specialized LLMs. In this paper, we study the vulnerability in web agents to malicious phrases embedded as HTML elements. To assess the security risks posed by this vulnerability, we present AgentBreaker, an IPI attack framework that autonomously composes adversarial phrases tailored to page-specific context. When processed by web agents, these DOM-embedded phrases induce adversarial behaviors, such as clicking attacker-designated HTML elements, posting attacker-provided text, and disclosing internal agent secrets. In our evaluation against five state-of-the-art web agents, AgentBreaker achieves an attack success rate of 71.7%–100% across 60 webpages sampled from Online-Mind2Web. We then propose practical defenses that not only mitigate observed threats but also address potential adaptive attacks. Our defenses reduce the attack success rate down to 1.7%. By conducting context-aware injection, AgentBreaker outperforms existing IPI frameworks, thereby accurately evaluating web agents’ susceptibility to IPI and providing stepping stones for countering this emerging threat.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- MUZZLE: Adaptive Agentic Red-Teaming of Web Agents Against Indirect Prompt Injection AttacksGeorgios Syros, Evan Rose, Brian Grinstead, Christoph Kerschbaumer 等USENIX Security 2026 · 被引用 18 次
- IPIGuard: A Novel Tool Dependency Graph-Based Defense Against Indirect Prompt Injection in LLM AgentsHengyu An, Jinghuai Zhang, Tianyu Du, Chunyi Zhou 等EMNLP 2025
- It's a TRAP! Task-Redirecting Agent Persuasion Benchmark for Web AgentsKarolina Korgul, Yushi Yang, Arkadiusz Drohomirecki, Piotr Blaszczyk 等ICML 2026 · 被引用 8 次
- A Large-scale Measurement of In-Page Prompt Injections Against LLM Web AgentsSoheil Khodayari, Xuenan Zhang, Bhupendra Acharya, Giancarlo PellegrinoCCS 2026
- ChatInject: Abusing Chat Templates for Prompt Injection in LLM AgentsHwan Chang, Yonghyun Jun, Hwanhee LeeICLR 2026 · 被引用 32 次
