Tracing the Dynamics of Refusal: Exploiting Latent Refusal Trajectories for Robust Jailbreak Detection
Xulin Hu, CHE WANG, Wei Yang Bryan Lim, Jianbo Gao, Zhong Chen
摘要
Representation Engineering analyses often characterize refusal using static directions extracted from terminal or pooled representations. We ask whether this view misses how refusal is constructed across layer-token positions. Using causal tracing, we identify a Refusal Trajectory: a sparse upstream activation pattern that often persists even when attacks such as GCG suppress terminal refusal signals. Based on this observation, we propose SALO (Sparse Activation Localization Operator), a lightweight white-box detector that operates on raw hidden-state volumes from a selected layer window. Across Qwen, Llama, and Mistral models, SALO improves jailbreak detection on several attack families under a fixed XSTest-calibrated operating point. We further analyze static RepE-style baselines, ROI sensitivity, adaptive GCG attacks, and encoded-input boundary cases, clarifying both the promise and limitations of refusal-trajectory monitoring.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper16
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 被引用 3,415 次
- Refusal in Language Models Is Mediated by a Single DirectionAndy Arditi, Oscar Obeso, Aaquib Syed, Daniel Paleka 等NeurIPS 2024 · 被引用 1,166 次
- Tree of Attacks: Jailbreaking Black-Box LLMs AutomaticallyAnay Mehrotra, Manolis Zampetakis, Paul Kassianik, Blaine Nelson 等NeurIPS 2024 · 被引用 835 次
相关 Paper
- Robust Harmful Features Under Jailbreak Attacks: Mechanistic Evidence from Attention Head Specialization in Large Language ModelsYanchen Yin, Dongqi Han, Linghui LiICML 2026 · 被引用 1 次
- Defending Jailbreak Attacks on Large Language Models via Manifold Trajectory KineticsHangtao Zhang, Yucheng Zhao, Sishun Liu, Ziqi Zhou 等USENIX Security 2026 · 被引用 3 次
- JailbreakScope: Interpreting Jailbreak Mechanism through Representation and Circuit AnalysesZeqing He, Zhibo Wang, Zhixuan Chu, Huiyu Xu 等USENIX Security 2026
- GraphShield: Graph-Theoretic Modeling of Network-Level Dynamics for Robust Jailbreak DetectionSunghee Dong, Sungwon Yi, Kangmin Bae, Jaeyoon Kim 等ICLR 2026
- SOM Directions Are Better than One: Multi-Directional Refusal Suppression in Language ModelsGiorgio Piras, Raffaele Mura, Fabio Brau, Luca Oneto 等AAAI 2026 · 被引用 4 次
