A Game-Theoretic Analysis of Attacks on Large Language Models via Compositional Skills
Xinbo Wu, Huan Zhang, Abhishek Umrawal, Lav Varshney
摘要
Warning: This paper contains potentially offensive or harmful text examples. As large language models grow increasingly capable, concerns about their safe deployment have intensified. While numerous alignment strategies aim to restrict harmful behavior, these defenses can still be circumvented through carefully designed adversarial prompts. In this work, we introduce a theoretical framework that formalizes a game between an attacker hiding its intent via compositional skills and a defender. Within this framework, we design a theoretical best-response attack strategy and show that it is closely related to many existing adversarial prompting methods. We further analyze the resulting game, characterize its equilibria, and reveal inherent advantages for the attacker. Drawing on our theoretical analysis, we also derive a provably optimal defense strategy. Empirically, we evaluate a practical instantiation of the theoretically optimal attack and observe stronger performance relative to existing adversarial prompting approaches in diverse settings encompassing different LLMs and benchmarks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper14
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust RefusalMantas Mazeika, Long Phan, Xuwang Yin, Andy Zou 等ICML 2024 · 被引用 1,031 次
- Tree of Attacks: Jailbreaking Black-Box LLMs AutomaticallyAnay Mehrotra, Manolis Zampetakis, Paul Kassianik, Blaine Nelson 等NeurIPS 2024 · 被引用 835 次
- AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language ModelsXiaogeng Liu, Nan Xu, Muhao Chen, Chaowei XiaoICLR 2024 · 被引用 722 次
相关 Paper
- SDD: Self-Degraded Defense against Malicious Fine-tuningZixuan Chen, Weikai Lu, Xin Lin, Ziqian ZengACL 2025
- Fundamental Limitations of Alignment in Large Language ModelsYotam Wolf, Noam Wies, Oshri Avnery, Yoav Levine 等ICML 2024 · 被引用 186 次
- Semantic Representation Attack against Aligned Large Language ModelsJiawei Lian, Jianhong Pan, Lefan Wang, Yi Wang 等NeurIPS 2025 · 被引用 3 次
- Defending Against Alignment-Breaking Attacks via Robustly Aligned LLMBochuan Cao, Yuanpu Cao, Lu Lin, Jinghui ChenACL 2024 · 被引用 34 次
- Safety Alignment of LMs via Non-cooperative GamesAnselm Paulus, Ilia Kulikov, Brandon Amos, REMI MUNOS 等ICML 2026 · 被引用 4 次
