Automatic Dialectic Jailbreak: A Framework for Generating Effective Jailbreak Strategies
Jianghai Yu, Yang Zhou, Zihan Zhou, Lingjuan Lyu, Da Yan, Ruoming Jin, Dejing Dou
摘要
Large language models (LLMs) can be jailbroken to produce malicious or unethical content with embedded jailbreaking prompts. Unfortunately, current jailbreak attack techniques suffer from adaptability issues due to reliance on the fixed evaluation models and incapability problems of surviving from a wide range of defense mechanisms. In this work, we propose to model the jailbreak attack problem as a Stackelberg multi-objective game between two LLMs engaged in a Hegelian-Dialectic-style debate enabling the automatic generation of jailbreak strategy (ADJ). In the ADJ, iterative thesis-antithesis-synthesis cycles of Hegelian dialectical reasoning are executed to guarantee that both attacker and defender can maximize their own utility while minimizing that of their opponent. We propose to map the optimization problem from the original parameter space into a Hilbert space via Haar wavelet transformation, for efficiently extracting localized and structurally significant information. In this functional space, we solve a convex multi-objective optimization problem to construct a common descent direction that better aligns with the objectives in the ADJ. In order to ensure sufficient descent for each objective in ADJ, we construct a subset of descent components and directly integrate them into the optimization objective. We theoretically validate the existence of a Pareto-Nash equilibrium achieved by our Automatic Dialectic Jailbreak method and demonstrate that our algorithm is able to converge to this Pareto-Nash equilibrium. Our source code is available at https://github. com/johnston1yu/ADJ Warning: This paper contains potentially harmful text.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper50
- HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust RefusalMantas Mazeika, Long Phan, Xuwang Yin, Andy Zou 等ICML 2024 · 被引用 1,031 次
- Tree of Attacks: Jailbreaking Black-Box LLMs AutomaticallyAnay Mehrotra, Manolis Zampetakis, Paul Kassianik, Blaine Nelson 等NeurIPS 2024 · 被引用 835 次
- AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language ModelsXiaogeng Liu, Nan Xu, Muhao Chen, Chaowei XiaoICLR 2024 · 被引用 722 次
- Catastrophic Jailbreak of Open-source LLMs via Exploiting GenerationYangsibo Huang, Samyak Gupta, Mengzhou Xia, Kai Li 等ICLR 2024 · 被引用 481 次
- Many-shot JailbreakingCem Anil, Esin Durmus, Nina Panickssery, Mrinank Sharma 等NeurIPS 2024 · 被引用 338 次
相关 Paper
- Efficient LLM-Jailbreaking via Multimodal-LLM JailbreakHaoxuan Ji, Zheng Lin, Zhenxing Niu, Xinbo Gao 等AAAI 2026 · 被引用 4 次
- Multi-Turn Jailbreaking Large Language Models via Attention ShiftingXiaohu Du, Fan Mo, Ming Wen, Tu Gu 等AAAI 2025 · 被引用 26 次
- from Benign import Toxic: Jailbreaking the Language Model via Adversarial MetaphorsYu Yan, Sheng Sun, Zenghao Duan, Teli Liu 等ACL 2025 · 被引用 14 次
- ASTRA: An Automated Framework for Strategy Discovery, Retrieval, and Evolution for Jailbreaking LLMsXu Liu, Yan Chen, Kan Ling, Yichi Zhu 等ACL 2026
- Towards Understanding Jailbreak Attacks in LLMs: A Representation Space AnalysisYuping Lin, Pengfei He, Han Xu, Yue Xing 等EMNLP 2024 · 被引用 6 次
