Obfuscated Activations Bypass LLM Latent-Space Defenses
Luke Bailey, Alex Serrano, Abhay Sheshadri, Mikhail Seleznyov, Jordan Taylor, Erik Jenner, Jacob Hilton, Stephen Casper, Carlos Guestrin, Scott Emmons
Abstract
Latent-space monitoring techniques have shown promise as defenses against LLM attacks. These defenses act as scanners to detect harmful activations before they lead to undesirable actions. This prompts the question: can models execute harmful behavior via inconspicuous latent states? Here, we study such obfuscated activations. Our results are nuanced. We show that state-of-the-art latent-space defenses---such as activation probes and latent OOD detection---are vulnerable to obfuscated activations. For example, against probes trained to classify harmfulness, our obfuscation attacks can reduce monitor recall from 100% down to 0% while still achieving a 90% jailbreaking success rate. However, we also find that certain probe architectures are more robust than others, and we discover the existence of an obfuscation tax: on a complex task (writing SQL code), evading monitors reduces model performance. Together, our results demonstrate white-box monitors are not robust to adversarial attack, while also providing concrete suggestions to alleviate, but not completely fix, this weakness.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 895dbdeb-8ef8-4cef-8e8b-e15cc95a1744Cited by top-tier papers8
- Detecting High-Stakes Interactions with Activation ProbesAlex McKenzie, Urja Pawar, Phil Blandfort, William Bankes et al.NeurIPS 2025 · 52 citations
- Language Models Are Capable of Metacognitive Monitoring and Control of Their Internal ActivationsJi-An Li, Huadong Xiong, Robert C. Wilson, Marcelo G. Mattar et al.NeurIPS 2025 · 49 citations
- Spilling the Beans: Teaching LLMs to Self-Report Their Hidden ObjectivesChloe Li, Mary Phuong, Daniel TanICLR 2026 · 14 citations
- Trojan-Speak: Bypassing Constitutional Classifiers with No Jailbreak Tax via Adversarial FinetuningBilgehan Sel, Xuanli He, Alwin Peng, Ming Jin et al.ICML 2026
- The Obfuscation Atlas: Mapping Where Honesty Emerges in RLVR with Deception ProbesMohammad Taufeeque, Stefan Heimersheim, Adam Gleave, Chris CundyICML 2026
Builds on30
- Jailbroken: How Does LLM Safety Training Fail?Alexander Wei, Nika Haghtalab, Jacob SteinhardtNeurIPS 2023 · 2,230 citations
- Refusal in Language Models Is Mediated by a Single DirectionAndy Arditi, Oscar Obeso, Aaquib Syed, Daniel Paleka et al.NeurIPS 2024 · 1,166 citations
- Sparse Autoencoders Find Highly Interpretable Features in Language ModelsRobert Huben, Hoagy Cunningham, Logan Riggs Smith, Aidan Ewart et al.ICLR 2024 · 1,072 citations
- Are aligned neural networks adversarially aligned?Nicholas Carlini, Milad Nasr, Christopher A. Choquette-Choo, Matthew Jagielski et al.NeurIPS 2023 · 412 citations
- Improving Alignment and Robustness with Circuit BreakersAndy Zou, Long Phan, Justin Wang, Derek Duenas et al.NeurIPS 2024 · 362 citations
Related papers
- Shaping the Safety Boundaries: Understanding and Defending Against Jailbreaks in Large Language ModelsLang Gao, Jiahui Geng, Xiangliang Zhang, Preslav Nakov et al.ACL 2025
- RobustKV: Defending Large Language Models against Jailbreak Attacks via KV EvictionTanqiu Jiang, Zian Wang, Jiacheng Liang, Changjiang Li et al.ICLR 2025
- Towards Understanding Jailbreak Attacks in LLMs: A Representation Space AnalysisYuping Lin, Pengfei He, Han Xu, Yue Xing et al.EMNLP 2024 · 6 citations
- MirrorShield: Towards Dynamic Adaptive Defense Against Jailbreaks via Entropy-Guided Mirror CraftingRui Pu, Chaozhuo Li, Rui Ha, Litian Zhang et al.AAAI 2026
- SpatialJB: How Text Distribution Art Becomes The "Jailbreak Key" for LLM GuardrailsZhiyi Mou, Jingyuan Yang, ZEHENG QIAN, Wangze Ni et al.ICML 2026 · 1 citation
