Obfuscated Activations Bypass LLM Latent-Space Defenses
Luke Bailey, Alex Serrano, Abhay Sheshadri, Mikhail Seleznyov, Jordan Taylor, Erik Jenner, Jacob Hilton, Stephen Casper, Carlos Guestrin, Scott Emmons
摘要
Latent-space monitoring techniques have shown promise as defenses against LLM attacks. These defenses act as scanners to detect harmful activations before they lead to undesirable actions. This prompts the question: can models execute harmful behavior via inconspicuous latent states? Here, we study such obfuscated activations. Our results are nuanced. We show that state-of-the-art latent-space defenses---such as activation probes and latent OOD detection---are vulnerable to obfuscated activations. For example, against probes trained to classify harmfulness, our obfuscation attacks can reduce monitor recall from 100% down to 0% while still achieving a 90% jailbreaking success rate. However, we also find that certain probe architectures are more robust than others, and we discover the existence of an obfuscation tax: on a complex task (writing SQL code), evading monitors reduces model performance. Together, our results demonstrate white-box monitors are not robust to adversarial attack, while also providing concrete suggestions to alleviate, but not completely fix, this weakness.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Detecting High-Stakes Interactions with Activation ProbesAlex McKenzie, Urja Pawar, Phil Blandfort, William Bankes 等NeurIPS 2025 · 被引用 52 次
- Language Models Are Capable of Metacognitive Monitoring and Control of Their Internal ActivationsJi-An Li, Huadong Xiong, Robert C. Wilson, Marcelo G. Mattar 等NeurIPS 2025 · 被引用 49 次
- Spilling the Beans: Teaching LLMs to Self-Report Their Hidden ObjectivesChloe Li, Mary Phuong, Daniel TanICLR 2026 · 被引用 14 次
- Trojan-Speak: Bypassing Constitutional Classifiers with No Jailbreak Tax via Adversarial FinetuningBilgehan Sel, Xuanli He, Alwin Peng, Ming Jin 等ICML 2026
- The Obfuscation Atlas: Mapping Where Honesty Emerges in RLVR with Deception ProbesMohammad Taufeeque, Stefan Heimersheim, Adam Gleave, Chris CundyICML 2026
它引用的顶会 Paper30
- Jailbroken: How Does LLM Safety Training Fail?Alexander Wei, Nika Haghtalab, Jacob SteinhardtNeurIPS 2023 · 被引用 2,230 次
- Refusal in Language Models Is Mediated by a Single DirectionAndy Arditi, Oscar Obeso, Aaquib Syed, Daniel Paleka 等NeurIPS 2024 · 被引用 1,166 次
- Sparse Autoencoders Find Highly Interpretable Features in Language ModelsRobert Huben, Hoagy Cunningham, Logan Riggs Smith, Aidan Ewart 等ICLR 2024 · 被引用 1,072 次
- Are aligned neural networks adversarially aligned?Nicholas Carlini, Milad Nasr, Christopher A. Choquette-Choo, Matthew Jagielski 等NeurIPS 2023 · 被引用 412 次
- Improving Alignment and Robustness with Circuit BreakersAndy Zou, Long Phan, Justin Wang, Derek Duenas 等NeurIPS 2024 · 被引用 362 次
相关 Paper
- Shaping the Safety Boundaries: Understanding and Defending Against Jailbreaks in Large Language ModelsLang Gao, Jiahui Geng, Xiangliang Zhang, Preslav Nakov 等ACL 2025
- RobustKV: Defending Large Language Models against Jailbreak Attacks via KV EvictionTanqiu Jiang, Zian Wang, Jiacheng Liang, Changjiang Li 等ICLR 2025
- Towards Understanding Jailbreak Attacks in LLMs: A Representation Space AnalysisYuping Lin, Pengfei He, Han Xu, Yue Xing 等EMNLP 2024 · 被引用 6 次
- MirrorShield: Towards Dynamic Adaptive Defense Against Jailbreaks via Entropy-Guided Mirror CraftingRui Pu, Chaozhuo Li, Rui Ha, Litian Zhang 等AAAI 2026
- SpatialJB: How Text Distribution Art Becomes The "Jailbreak Key" for LLM GuardrailsZhiyi Mou, Jingyuan Yang, ZEHENG QIAN, Wangze Ni 等ICML 2026 · 被引用 1 次
