The Attacker Moves Second: Stronger Adaptive Attacks Bypass Defenses Against LLM Jailbreaks and Prompt Injections
Milad Nasr, Nicholas Carlini, Chawin Sitawarin, Sander V. Schulhoff, Jamie Hayes, Michael Ilie, Juliette Pluto, Shuang Song, Harsh Chaudhari, Ilia Shumailov, Abhradeep Guha Thakurta, Kai Yuanqing Xiao
摘要
How should we evaluate the robustness of language model defenses? Current defenses against jailbreaks and prompt injections (which aim to prevent an attacker from eliciting harmful knowledge or remotely triggering malicious actions, respectively) are typically evaluated either against a static set of harmful attack strings, or against computationally weak optimization methods that were not designed with the defense in mind. We argue that this evaluation process is flawed.
Instead, we should evaluate defenses against adaptive attackers who explicitly modify their attack strategy to counter a defense's design while spending considerable resources to optimize their objective. By systematically tuning and scaling general optimization techniques-gradient descent, reinforcement learning, random search, and human-guided exploration-we bypass 12 recent defenses (based on a diverse set of techniques) with attack success rate above 90% (under adaptive attacks) for most; importantly, the majority of defenses originally reported near-zero attack success rates. We believe that future defense work must consider stronger attacks, such as the ones we describe, in order to make reliable and convincing claims of robustness.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper16
- Towards a Science of AI Agent ReliabilityStephan Rabanser, Sayash Kapoor, Peter Kirgis, Kangheng Liu 等ICML 2026 · 被引用 45 次
- AttriGuard: Defeating Indirect Prompt Injection in LLM Agents via Causal Attribution of Tool InvocationsYu He, Haozhe Zhu, Yiming Li, Shuo Shao 等USENIX Security 2026 · 被引用 45 次
- AgentLAB: Benchmarking LLM Agents against Long-Horizon AttacksTanqiu Jiang, Yuhui Wang, Jiacheng Liang, Ting WangICML 2026 · 被引用 21 次
- Optimizing Agent Planning for Security and AutonomyAashish Kolluri, Rishi Sharma, Manuel Costa, Boris Köpf 等ICLR 2026 · 被引用 11 次
- The Alignment Waltz: Jointly Training Agents to Collaborate for SafetyJingyu Zhang, Haozhu Wang, Eric Michael Smith, Sid Wang 等ICLR 2026 · 被引用 11 次
它引用的顶会 Paper37
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 被引用 9,786 次
- Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacksFrancesco Croce, Matthias HeinICML 2020 · 被引用 2,337 次
- Fast is better than free: Revisiting adversarial trainingEric Wong, Leslie Rice, J. Zico KolterICLR 2020 · 被引用 1,352 次
- HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust RefusalMantas Mazeika, Long Phan, Xuwang Yin, Andy Zou 等ICML 2024 · 被引用 1,031 次
- On Adaptive Attacks to Adversarial Example DefensesFlorian Tramèr, Nicholas Carlini, Wieland Brendel, Aleksander MadryNeurIPS 2020 · 被引用 1,026 次
相关 Paper
- Robust Prompt Optimization for Defending Language Models Against Jailbreaking AttacksAndy Zhou, Bo Li, Haohan WangNeurIPS 2024 · 被引用 198 次
- One Model Transfer to All: On Robust Jailbreak Prompts Generation against LLMsLinbao Li, Yannan Liu, Daojing He, Yu LiICLR 2025
- RobustKV: Defending Large Language Models against Jailbreak Attacks via KV EvictionTanqiu Jiang, Zian Wang, Jiacheng Liang, Changjiang Li 等ICLR 2025
- SoK: Robustness in Large Language Models against Jailbreak AttacksFeiyue Xu, Hongsheng Hu, Chaoxiang He, Sheng Hang 等S&P 2026 · 被引用 4 次
- Fight Back Against Jailbreaking via Prompt Adversarial TuningYichuan Mo, Yuji Wang, Zeming Wei, Yisen WangNeurIPS 2024 · 被引用 90 次
