JPU: Bridging Jailbreak Defense and Unlearning via On-Policy Path Rectification
Xi Wang, Songlei Jian, Shasha Li, Xiaopeng Li, Zhaoye Li, Bin Ji, Baosheng Wang, Jie Yu
摘要
Despite extensive safety alignment, Large Language Models (LLMs) often fail against jailbreak attacks. While machine unlearning has emerged as a promising defense by erasing specific harmful parameters, current methods remain vulnerable to diverse jailbreaks. We first conduct an empirical study and discover that this failure mechanism is caused by jailbreaks primarily activating non-erased parameters in the intermediate layers. Further, by probing the underlying mechanism through which these circumvented parameters reassemble into the prohibited output, we verify the persistent existence of dynamic and show that the inability to rectify them constitutes the fundamental gap in existing unlearning defenses. To bridge this gap, we propose ailbreak ath nlearning (JPU), which is the first to rectify dynamic jailbreak paths towards safety anchors by dynamically mining on-policy adversarial samples to expose vulnerabilities and identify jailbreak paths. Extensive experiments demonstrate that JPU significantly enhances jailbreak resistance against dynamic attacks while preserving the model's utility.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Rethinking Jailbreak Detection of Large Vision Language Models with Representational Contrastive ScoringPeichun Hua, Hao Li, Shanghao Shi, Zhiyuan Yu 等ACL 2026 · 被引用 8 次
- TROJail: Trajectory-Level Optimization for Multi-Turn Large Language Model Jailbreaks with Process RewardsXiqiao Xiong, Ouxiang Li, Zhuo Liu, Moxin Li 等ACL 2026 · 被引用 7 次
它引用的顶会 Paper11
- AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language ModelsXiaogeng Liu, Nan Xu, Muhao Chen, Chaowei XiaoICLR 2024 · 被引用 722 次
- Catastrophic Jailbreak of Open-source LLMs via Exploiting GenerationYangsibo Huang, Samyak Gupta, Mengzhou Xia, Kai Li 等ICLR 2024 · 被引用 481 次
- Improving Alignment and Robustness with Circuit BreakersAndy Zou, Long Phan, Justin Wang, Derek Duenas 等NeurIPS 2024 · 被引用 362 次
- Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank ModificationsBoyi Wei, Kaixuan Huang, Yangsibo Huang, Tinghao Xie 等ICML 2024 · 被引用 215 次
- RAIN: Your Language Models Can Align Themselves without FinetuningYuhui Li, Fangyun Wei, Jinjing Zhao, Chao Zhang 等ICLR 2024 · 被引用 171 次
相关 Paper
- Dynamic Deep Prompt Optimization for Defending Against Jailbreak Attacks on LLMsDoniyorkhon Obidov, Honggang Yu, Xiaolong Guo, Kaichen YangAAAI 2026
- Safety Alignment via Constrained Knowledge UnlearningZesheng Shi, Yucheng Zhou, Jing Li, Yuxin Jin 等ACL 2025 · 被引用 8 次
- JULI: Jailbreak Large Language Models by Self-IntrospectionZhixian Wang, Zhanhao Hu, David A. WagnerICLR 2026 · 被引用 3 次
- Weak-to-Strong Jailbreaking on Large Language ModelsXuandong Zhao, Xianjun Yang, Tianyu Pang, Chao Du 等ICML 2025
- Immune: Improving Safety Against Jailbreaks in Multi-modal LLMs via Inference-Time AlignmentSoumya Suvra Ghosal, Souradip Chakraborty, Vaibhav Singh, Tianrui Guan 等CVPR 2025
