Fine-grained Pluggable Gradient Ascent for Knowledge Unlearning in Language Models
Xiaohua Feng, Chaochao Chen, Yuyuan Li, Zibin Lin
摘要
Pre-trained language models acquire knowledge from vast amounts of text data, which can inadvertently contain sensitive information. To mitigate the presence of undesirable knowledge, the task of knowledge unlearning becomes crucial for language models. Previous research relies on gradient ascent methods to achieve knowledge unlearning, which is simple and effective. However, this approach calculates all the gradients of tokens in the sequence, potentially compromising the general ability of language models. To overcome this limitation, we propose an adaptive objective that calculates gradients with fine-grained control specifically targeting sensitive tokens. Our adaptive objective is pluggable, ensuring simplicity and enabling extension to the regularization-based framework that utilizes non-target data or other models to preserve general ability. Through extensive experiments targeting the removal of typical sensitive data, we demonstrate that our proposed method enhances the general ability of language models while achieving knowledge unlearning. Additionally, it demonstrates the capability to adapt to behavior alignment, eliminating all the undesirable knowledge within a specific domain.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Keeping an Eye on LLM Unlearning: The Hidden Risk and RemedyJie Ren, Zhenwei Dai, Xianfeng Tang, Yue Xing 等NeurIPS 2025 · 被引用 11 次
- Leveraging Machine Unlearning for Cost-Efficient Preference AlignmentXiaoHua Feng, Yuyuan Li, HuWei Ji, Li Zhang 等ICML 2026 · 被引用 4 次
- FedAU2: Attribute Unlearning for User-Level Federated Recommender Systems with Adaptive and Robust Adversarial TrainingYuyuan Li, Junjie Fang, Fengyuan Yu, Xichun Sheng 等AAAI 2026 · 被引用 1 次
- Unbiased Rectification for Sequential Recommender Systems Under Fake OrdersQiyu Qin, Yichen Li, Haozhao Wang, Cheng Wang 等AAAI 2026
- CAP: Controllable Alignment Prompting for Unlearning in LLMsZhaokun Wang, Jinyu Guo, Jingwen Pu, Hongli Pu 等ACL 2026
它引用的顶会 Paper23
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao 等AAAI 2020 · 被引用 2,916 次
- Extracting Training Data from Large Language ModelsNicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski 等USENIX Security 2021 · 被引用 2,866 次
- Machine UnlearningLucas Bourtoule, Varun Chandrasekaran, Christopher A. Choquette-Choo, Hengrui Jia 等S&P 2021 · 被引用 1,381 次
- Certified Data Removal from Machine Learning ModelsChuan Guo, Tom Goldstein, Awni Y. Hannun, Laurens van der MaatenICML 2020 · 被引用 633 次
相关 Paper
- Knowledge Unlearning for Mitigating Privacy Risks in Language ModelsJoel Jang, Dongkeun Yoon, Sohee Yang, Sungmin Cha 等ACL 2023 · 被引用 48 次
- Explainable LLM Unlearning through ReasoningJunfeng Liao, Qizhou Wang, Shanshan Ye, Xin Yu 等ICLR 2026 · 被引用 8 次
- LLM-Eraser: Optimizing Large Language Model Unlearning through Selective PruningShengming Zhang, Le Zhang, Jingbo Zhou, Zhi Zheng 等KDD 2025 · 被引用 2 次
- Erase or Hide? Suppressing Spurious Unlearning Neurons for Robust UnlearningNakyeong Yang, Dong-Kyum Kim, Jea Kwon, Minsung Kim 等ICLR 2026 · 被引用 7 次
- Decoding-Unlearning: Fact Forgetting via Entropy-Guided InferenceJingwen Pu, Mingjun Shi, Xinrui Ren, Yizhe Wang 等ACL 2026
