It's a TRAP! Task-Redirecting Agent Persuasion Benchmark for Web Agents
Karolina Korgul, Yushi Yang, Arkadiusz Drohomirecki, Piotr Blaszczyk, Will Howard, Lukas Aichberger, Chris Russell, Phil Torr, Adam Mahdi, Adel Bibi
摘要
Web-based agents powered by large language models are increasingly used for tasks such as email management or professional networking. Their reliance on dynamic web content, however, makes them vulnerable to prompt injection attacks: adversarial instructions hidden in interface elements that persuade the agent to divert from its original task. We introduce the Task-Redirecting Agent Persuasion Benchmark (TRAP), a benchmark for studying how persuasion techniques misguide autonomous web agents on realistic tasks. Across six frontier models, agents are susceptible to prompt injection in 25% of tasks on average (13% for GPT-5 to 43% for DeepSeek-R1), with small interface or contextual changes often doubling success rates and revealing systemic, psychologically driven vulnerabilities in web-based agents. We also provide a modular social-engineering injection framework with controlled experiments on high-fidelity website clones, allowing for further benchmark expansion.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper3
- LLM Evaluators Recognize and Favor Their Own GenerationsArjun Panickssery, Samuel R. Bowman, Shi FengNeurIPS 2024 · 被引用 865 次
- SEC-bench: Automated Benchmarking of LLM Agents on Real-World Software Security TasksHwiwon Lee, Ziqi Zhang, Hanxiao Lu, Lingming ZhangNeurIPS 2025 · 被引用 86 次
- AgentHarm: A Benchmark for Measuring Harmfulness of LLM AgentsMaksym Andriushchenko, Alexandra Souly, Mateusz Dziemian, Derek Duenas 等ICLR 2025 · 被引用 4 次
相关 Paper
- ChatInject: Abusing Chat Templates for Prompt Injection in LLM AgentsHwan Chang, Yonghyun Jun, Hwanhee LeeICLR 2026 · 被引用 32 次
- TRAP: Targeted Redirecting of Agentic PreferencesHangoo Kang, Jehyeok Yeon, Gagandeep SinghNeurIPS 2025 · 被引用 7 次
- AgentBreaker: Evaluating Context-Aware Indirect Prompt Injection Risks in Modern Web AgentsYongbi Son, Changoo Lee, Dongwon Shin, Byoungyoung Lee 等ISSTA 2026
- WebInject: Prompt Injection Attack to Web AgentsXilong Wang, John Bloch, Zedian Shao, Yuepeng Hu 等EMNLP 2025 · 被引用 1 次
- When Efficiency Becomes a Vulnerability: Computational Cost Attacks on WebAgentsLiang-Bo Ning, Yuchen Zhu, Heqing Huang, Xin Wang 等ACL 2026
