Dual Intention Escape: Penetrating and Toxic Jailbreak Attack against Large Language Models
Yanni Xue, Jiakai Wang, Zixin Yin, Yuqing Ma, Haotong Qin, Renshuai Tao, Xianglong Liu
摘要
Recently, the jailbreak attack, which generates adversarial prompts to bypass safety measures and mislead large language models (LLMs) to output harmful answers, has attracted extensive interest due to its potential to reveal the vulnerabilities of LLMs. However, ignoring the exploitation of the characteristics in intention understanding, existing studies could only generate prompts with weak attacking ability, failing to evade defenses (e.g., sensitive word detect) and causing malice(e.g., harmful outputs). Motivated by the mechanism in the psychology of human misjudgment, we propose a dual intention escape (DIE) jailbreak attack framework to generate more stealthy and toxic prompts to deceive LLMs to output harmful content. For stealthiness, inspired by the anchoring effect, we designed the Intention-anchored Malicious Concealment(IMC) module that hides the harmful intention behind a generated anchor intention by the recursive decomposition block and contrary intention nesting block. Since the anchor intention will be received first, the LLMs might pay less attention to the harmful intention and enter response status. For toxicity, we propose the Intention-reinforced Malicious Inducement (IMI) module based on the availability bias mechanism in a progressive malicious prompting approach. Due to the ongoing emergence of statements correlated to harmful intentions, the output content of LLMs will be closer to these more accessible intentions, i.e., more toxic. We conducted extensive experiments under black-box settings, supporting that DIE could achieve 100% ASR-R and 92.9% ASR-G against GPT3.5-turbo.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper3
- SoK: Robustness in Large Language Models against Jailbreak AttacksFeiyue Xu, Hongsheng Hu, Chaoxiang He, Sheng Hang 等S&P 2026 · 被引用 4 次
- Has the Two-Decade-Old Prophecy Come True? Artificial Bad Intelligence Triggered by Merely a Single-Bit Flip in Large Language ModelsYu Yan, Siqi Lu, Yang Gao, Zhaoxuan Li 等WWW 2026 · 被引用 1 次
- Exploiting Synergistic Cognitive Biases to Bypass Safety in LLMsXikang Yang, Biyu Zhou, Xuehai Tang, Jizhong Han 等AAAI 2026
相关 Paper
- Cannot See the Forest for the Trees: Invoking Heuristics and Biases to Elicit Irrational Choices of LLMsHaoming Yang, Ke Ma, Xiaojun Jia, Yingfei Sun 等ICML 2025
- Towards Understanding Jailbreak Attacks in LLMs: A Representation Space AnalysisYuping Lin, Pengfei He, Han Xu, Yue Xing 等EMNLP 2024 · 被引用 6 次
- Multi-Turn Jailbreaking Large Language Models via Attention ShiftingXiaohu Du, Fan Mo, Ming Wen, Tu Gu 等AAAI 2025 · 被引用 26 次
- Large Language Models Are Involuntary Truth-Tellers: Exploiting Fallacy Failure for Jailbreak AttacksYue Zhou, Henry Peng Zou, Barbara Di Eugenio, Yang ZhangEMNLP 2024 · 被引用 3 次
- from Benign import Toxic: Jailbreaking the Language Model via Adversarial MetaphorsYu Yan, Sheng Sun, Zenghao Duan, Teli Liu 等ACL 2025 · 被引用 14 次
