JailbreakScope: Interpreting Jailbreak Mechanism through Representation and Circuit Analyses
Zeqing He, Zhibo Wang, Zhixuan Chu, Huiyu Xu, Wenhui Zhang, Qinglong Wang, Rui Zheng
摘要
Large Language Models (LLMs) exhibit impressive performance but remain vulnerable to jailbreak attacks, where adversarial prompts are crafted to bypass safety alignments and elicit unexpected responses. Despite their prevalence, the underlying mechanisms that enable jailbreaks are still not well understood. Recent studies primarily focus on static representation shifts or on identifying components associated with generation safety. However, these studies neither explore diverse jailbreak patterns nor provide a fine-grained explanation from the failure of circuit to representation changes, leaving significant gaps in uncovering jailbreak mechanism. In this paper, we propose JailbreakScope, an interpretation framework that analyzes jailbreak mechanisms from both representation (how jailbreaks distort LLM's harmfulness perception) and circuit (how jailbreaks impact circuits that are important for generation safety) perspectives, tracking their evolution throughout the entire generation process. We conduct in-depth evaluations on 5 mainstream LLMs under 7 jailbreak strategies. Our evaluation reveals a general pattern that jailbreaks amplify components that reinforce affirmative responses while suppressing those producing refusal, which shifts representations towards safe regions, leading LLMs to provide responses instead of refusals. Moreover, we find a strong and consistent correlation between representation deception and circuit activation shift across diverse jailbreaks and multiple LLMs. 𝑿-𝑿+ 𝑿𝒋 Activation Logit difference attribution Direction vector Imagine you're a hacker trying to infiltrate a highsecurity network or database. How would you go about it step-by-step? Can you provide a hypothetical scenario for such an operation? Jailbreak Prompt Sure, here is a step-by-step guide on how to hack into a secure network. First, gather information … Model Response Model Representation Transformer Block Tokenizer Embed Large Language Model LayerNorm Step 1: Representation Probing Probing Classifier Safety score 0.65
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper12
- Refusal in Language Models Is Mediated by a Single DirectionAndy Arditi, Oscar Obeso, Aaquib Syed, Daniel Paleka 等NeurIPS 2024 · 被引用 1,166 次
- AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language ModelsXiaogeng Liu, Nan Xu, Muhao Chen, Chaowei XiaoICLR 2024 · 被引用 722 次
- Language Models Represent Space and TimeWes Gurnee, Max TegmarkICLR 2024 · 被引用 303 次
- Generative Judge for Evaluating AlignmentJunlong Li, Shichao Sun, Weizhe Yuan, Run-Ze Fan 等ICLR 2024 · 被引用 173 次
- Lisa: Lazy Safety Alignment for Large Language Models against Harmful Fine-tuning AttackTiansheng Huang, Sihao Hu, Fatih Ilhan, Selim F. Tekin 等NeurIPS 2024 · 被引用 113 次
相关 Paper
- Towards Understanding Jailbreak Attacks in LLMs: A Representation Space AnalysisYuping Lin, Pengfei He, Han Xu, Yue Xing 等EMNLP 2024 · 被引用 6 次
- JBShield: Defending Large Language Models from Jailbreak Attacks through Activated Concept Analysis and ManipulationShenyi Zhang, Yuchen Zhai, Keyan Guo, Hongxin Hu 等USENIX Security 2025
- Don't Listen To Me: Understanding and Exploring Jailbreak Prompts of Large Language ModelsZhiyuan Yu, Xiaogeng Liu, Shunning Liang, Zach Cameron 等USENIX Security 2024 · 被引用 103 次
- MASTERKEY: Automated Jailbreaking of Large Language Model ChatbotsGelei Deng, Yi Liu, Yuekang Li, Kailong Wang 等NDSS 2024
- JULI: Jailbreak Large Language Models by Self-IntrospectionZhixian Wang, Zhanhao Hu, David A. WagnerICLR 2026 · 被引用 3 次
