Lune

ICML2026顶会

Steering Beyond the Support: Adversarial Training on Unsupervised Jailbroken Activation Simulation

LUOYU CHEN, Weiqi Wang, Zhiyi Tian, Chenhan Zhang, Feng Wu, Jianhuan Huang, Ahmed Asiri, Shui Yu

2026年份

摘要

Jailbreak prompts can trigger harmful completions from aligned LLMs. Accordingly, safety steering has been proposed as a family of test-time activation interventions that redirect jailbreak activations toward refusal while preserving benign utility. However, existing steering methods are fundamentally supervised and tied to a static, limited training set, whereas real jailbreaks evolve and are often out of distribution relative to the training set, leading to failures on unseen attacks. In this paper, we tackle this failure by developing a zero-shot defense. Based on unsupervised latent direction discovery, we directly simulate jailbroken activations without any knowledge of jailbreak strategy. To build a defense mechanism on top of this simulation, we propose a bi-level adversarial training framework. In the inner step, we simulate diverse jailbroken activations by extrapolating from refusal-state harmful-request activations via unsupervised latent direction discovery. In the outer step, we train a potential-induced steering field to push these adversarial jailbroken states into refusal regions while leaving benign activations unchanged. Across three LLMs and six classical jailbreak families, our method achieves strong defense with attack success rates mostly below 5%. We further analyze the increasing subspace coverage of our simulated jailbroken activations over real jailbreaks throughout training, which helps explain the increasing robustness of our defense mechanism.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

lune papers fulltext 235dbbfa-1380-4f81-9a1b-2f4469fcd4cf

它引用的顶会 Paper14

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖