Steering Beyond the Support: Adversarial Training on Unsupervised Jailbroken Activation Simulation
LUOYU CHEN, Weiqi Wang, Zhiyi Tian, Chenhan Zhang, Feng Wu, Jianhuan Huang, Ahmed Asiri, Shui Yu
Abstract
Jailbreak prompts can trigger harmful completions from aligned LLMs. Accordingly, safety steering has been proposed as a family of test-time activation interventions that redirect jailbreak activations toward refusal while preserving benign utility. However, existing steering methods are fundamentally supervised and tied to a static, limited training set, whereas real jailbreaks evolve and are often out of distribution relative to the training set, leading to failures on unseen attacks. In this paper, we tackle this failure by developing a zero-shot defense. Based on unsupervised latent direction discovery, we directly simulate jailbroken activations without any knowledge of jailbreak strategy. To build a defense mechanism on top of this simulation, we propose a bi-level adversarial training framework. In the inner step, we simulate diverse jailbroken activations by extrapolating from refusal-state harmful-request activations via unsupervised latent direction discovery. In the outer step, we train a potential-induced steering field to push these adversarial jailbroken states into refusal regions while leaving benign activations unchanged. Across three LLMs and six classical jailbreak families, our method achieves strong defense with attack success rates mostly below 5%. We further analyze the increasing subspace coverage of our simulated jailbroken activations over real jailbreaks throughout training, which helps explain the increasing robustness of our defense mechanism.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 235dbbfa-1380-4f81-9a1b-2f4469fcd4cfBuilds on14
- Refusal in Language Models Is Mediated by a Single DirectionAndy Arditi, Oscar Obeso, Aaquib Syed, Daniel Paleka et al.NeurIPS 2024 · 1,166 citations
- Sparse Autoencoders Find Highly Interpretable Features in Language ModelsRobert Huben, Hoagy Cunningham, Logan Riggs Smith, Aidan Ewart et al.ICLR 2024 · 1,072 citations
- HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust RefusalMantas Mazeika, Long Phan, Xuwang Yin, Andy Zou et al.ICML 2024 · 1,031 citations
- Tree of Attacks: Jailbreaking Black-Box LLMs AutomaticallyAnay Mehrotra, Manolis Zampetakis, Paul Kassianik, Blaine Nelson et al.NeurIPS 2024 · 835 citations
- Improving Alignment and Robustness with Circuit BreakersAndy Zou, Long Phan, Justin Wang, Derek Duenas et al.NeurIPS 2024 · 362 citations
Related papers
- InDe-LLM: Defending against Jailbreak Attacks in LLM-Powered Systems via Intention DisentanglingYujue Wang, Quan Zhang, Chijin Zhou, Gwihwan Go et al.FSE 2026
- AlphaSteer: Learning Refusal Steering with Principled Null-Space ConstraintLeheng Sheng, Changshuo Shen, Weixiang Zhao, Junfeng Fang et al.ICLR 2026 · 52 citations
- Principled Steering via Null-space Projection for Jailbreak Defense in Vision-Language ModelsXingyu Zhu, Beier Zhu, Shuo Wang, Junfeng Fang et al.CVPR 2026 · 5 citations
- AdaSteer: Your Aligned LLM is Inherently an Adaptive Jailbreak DefenderWeixiang Zhao, Jiahe Guo, Yulin Hu, Yang Deng et al.EMNLP 2025
- Lifelong Safety Alignment for Language ModelsHaoyu Wang, Yifei Zhao, Zeyu Qin, Chao Du et al.NeurIPS 2025 · 18 citations
