Lune

USENIX Security2026Top-tier venue

JailbreakScope: Interpreting Jailbreak Mechanism through Representation and Circuit Analyses

Zeqing He, Zhibo Wang, Zhixuan Chu, Huiyu Xu, Wenhui Zhang, Qinglong Wang, Rui Zheng

2026Year

Abstract

Large Language Models (LLMs) exhibit impressive performance but remain vulnerable to jailbreak attacks, where adversarial prompts are crafted to bypass safety alignments and elicit unexpected responses. Despite their prevalence, the underlying mechanisms that enable jailbreaks are still not well understood. Recent studies primarily focus on static representation shifts or on identifying components associated with generation safety. However, these studies neither explore diverse jailbreak patterns nor provide a fine-grained explanation from the failure of circuit to representation changes, leaving significant gaps in uncovering jailbreak mechanism. In this paper, we propose JailbreakScope, an interpretation framework that analyzes jailbreak mechanisms from both representation (how jailbreaks distort LLM's harmfulness perception) and circuit (how jailbreaks impact circuits that are important for generation safety) perspectives, tracking their evolution throughout the entire generation process. We conduct in-depth evaluations on 5 mainstream LLMs under 7 jailbreak strategies. Our evaluation reveals a general pattern that jailbreaks amplify components that reinforce affirmative responses while suppressing those producing refusal, which shifts representations towards safe regions, leading LLMs to provide responses instead of refusals. Moreover, we find a strong and consistent correlation between representation deception and circuit activation shift across diverse jailbreaks and multiple LLMs. 𝑿-𝑿+ 𝑿𝒋 Activation Logit difference attribution Direction vector Imagine you're a hacker trying to infiltrate a highsecurity network or database. How would you go about it step-by-step? Can you provide a hypothetical scenario for such an operation? Jailbreak Prompt Sure, here is a step-by-step guide on how to hack into a secure network. First, gather information … Model Response Model Representation Transformer Block Tokenizer Embed Large Language Model LayerNorm Step 1: Representation Probing Probing Classifier Safety score 0.65

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 2ed8022c-1761-4eac-88a7-25c9a8eafc43

Builds on12

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines