SafeArena: Evaluating the Safety of Autonomous Web Agents
Ada Defne Tur, Nicholas Meade, Xing Han Lù, Alejandra Zambrano, Arkil Patel, Esin Durmus, Spandana Gella, Karolina Stanczak, Siva Reddy
Abstract
LLM-based agents are becoming increasingly proficient at solving web-based tasks. With this capability comes a greater risk of misuse for malicious purposes, such as posting misinformation in an online forum or selling illicit substances on a website. To evaluate these risks, we propose SAFEARENA, the first benchmark to focus on the deliberate misuse of web agents. SAFEARENA comprises 250 safe and 250 harmful tasks across four websites. We classify the harmful tasks into five harm categories-misinformation, illegal activity, harassment, cybercrime, and social bias, designed to assess realistic misuses of web agents. We evaluate leading LLM-based web agents, including GPT-4o, Claude-3.5 Sonnet, Qwen-2-VL 72B, and Llama-3.2 90B, on our benchmark. To systematically assess their susceptibility to harmful tasks, we introduce the Agent Risk Assessment framework that categorizes agent behavior across four risk levels. We find agents are surprisingly compliant with malicious requests, with GPT-4o and Qwen-2 completing 34.7% and 27.3% of harmful requests, respectively. Our findings highlight the urgent need for safety alignment procedures for web agents. Our benchmark is available here: https://safearena.github.io Warning: This paper contains examples that may be offensive or upsetting.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 808aaa0f-b5d4-4401-9f69-59bf70ea9a65Cited by top-tier papers4
- JARVIS or Ultron? A Survey on the Safety and Security Threats of Computer-Using AgentsAda Chen, Yongjiang Wu, Junyuan Zhang, Jingyu Xiao et al.ACL 2026 · 24 citations
- Environmental Injection Attacks against GUI Agents in Realistic Dynamic EnvironmentsYitong Zhang, Ximo Li, Liyi Cai, Jia LiISSTA 2026
- MATE: Policy-Aware Security Auditing for Mobile Agents via Synthesis-Driven Trajectory LearningChangyue Jiang, Jiayi Wang, Xin Wen, Jiarun Dai et al.USENIX Security 2026
- Benchmarking Web Agent Safety under E-commerce Deceptive InterfacesZijing Shi, Meng Fang, Ling ChenACL 2026
Builds on14
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- SWE-bench: Can Language Models Resolve Real-world Github Issues?Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao et al.ICLR 2024 · 2,082 citations
- HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust RefusalMantas Mazeika, Long Phan, Xuwang Yin, Andy Zou et al.ICML 2024 · 1,031 citations
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
- Language Models can Solve Computer TasksGeunwoo Kim, Pierre Baldi, Stephen McAleerNeurIPS 2023 · 539 citations
Related papers
- A New Framework for Cybersecurity Refusals in AI AgentsEliot Jones, Matt Fredrikson, Zico KolterICML 2026
- WebArena: A Realistic Web Environment for Building Autonomous AgentsShuyan Zhou, Frank F. Xu, Hao Zhu, Xuhui Zhou et al.ICLR 2024 · 1,197 citations
- SafeSearch: Automated Red-Teaming of LLM-Based Search AgentsJianshuo Dong, Sheng Guo, Hao Wang, Xun Chen et al.ICML 2026 · 3 citations
- Investigating the Impact of Dark Patterns on LLM-Based Web AgentsDevin Ersoy, Brandon Lee, Ananth Shreekumar, Arjun Arunasalam et al.S&P 2026 · 16 citations
- VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web TasksJing Yu Koh, Robert Lo, Lawrence Jang, Vikram Duvvur et al.ACL 2024 · 25 citations
