Safety Alignment via Constrained Knowledge Unlearning
Zesheng Shi, Yucheng Zhou, Jing Li, Yuxin Jin, Yu Li, Daojing He, Fangming Liu, Saleh Alharbi, Jun Yu, Min Zhang
Abstract
Despite significant progress in safety alignment, large language models (LLMs) remain susceptible to jailbreak attacks. Existing defense mechanisms have not fully deleted harmful knowledge in LLMs, which allows such attacks to bypass safeguards and produce harmful outputs. To address this challenge, we propose a novel safety alignment strategy, Constrained Knowledge Unlearning (CKU), which focuses on two primary objectives: knowledge localization and retention, and unlearning harmful knowledge. CKU works by scoring neurons in specific multilayer perceptron (MLP) layers to identify a subset U of neurons associated with useful knowledge. During the unlearning process, CKU prunes the gradients of neurons in U to preserve valuable knowledge while effectively mitigating harmful content. Experimental results demonstrate that CKU significantly enhances model safety without compromising overall performance, offering a superior balance between safety and utility compared to existing methods. Additionally, our analysis of neuron knowledge sensitivity across various MLP layers provides valuable insights into the mechanics of safety alignment and model knowledge editing. This paper contains harmful data and modelgenerated content that may be offensive.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 52afadca-b08a-4d81-89e9-e1158e27deffCited by top-tier papers4
- Multi-objective Large Language Model Alignment with Hierarchical ExpertsZhuo Li, Guodong DU, Weiyang Guo, Yigeng Zhou et al.ICLR 2026 · 17 citations
- Rethinking Jailbreak Detection of Large Vision Language Models with Representational Contrastive ScoringPeichun Hua, Hao Li, Shanghao Shi, Zhiyuan Yu et al.ACL 2026 · 8 citations
- JPU: Bridging Jailbreak Defense and Unlearning via On-Policy Path RectificationXi Wang, Songlei Jian, Shasha Li, Xiaopeng Li et al.ACL 2026 · 2 citations
- Team-Based Self-Play With Dual Adaptive Weighting for Fine-Tuning LLMsWu Li, Yigeng Zhou, Zesheng Shi, Yequan Wang et al.ACL 2026
Builds on30
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu et al.ICLR 2022 · 4,966 citations
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 3,037 citations
- Machine UnlearningLucas Bourtoule, Varun Chandrasekaran, Christopher A. Choquette-Choo, Hengrui Jia et al.S&P 2021 · 1,381 citations
Related papers
- From "Sure" to "Sorry": Detecting Jailbreak in Large Vision Language Model via JailNeuronsYuyou Gan, Qingming Li, Junhao Li, Zhi Chen et al.ICLR 2026
- SABER: Uncovering Vulnerabilities in Safety Alignment via Cross-Layer Residual ConnectionMaithili Joshi, Palash Nandi, Tanmoy ChakrabortyEMNLP 2025 · 1 citation
- Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank ModificationsBoyi Wei, Kaixuan Huang, Yangsibo Huang, Tinghao Xie et al.ICML 2024 · 215 citations
- Improving LLM Safety Alignment with Dual-Objective OptimizationXuandong Zhao, Will Cai, Tianneng Shi, David Huang et al.ICML 2025
- Alignment-Enhanced Decoding: Defending Jailbreaks via Token-Level Adaptive Refining of Probability DistributionsQuan Liu, Zhenhong Zhou, Longzhu He, Yi Liu et al.EMNLP 2024 · 1 citation
