Rethinking LLM Unlearning Objectives: A Gradient Perspective and Go Beyond
Qizhou Wang, Jin Peng Zhou, Zhanke Zhou, Saebyeol Shin, Bo Han, Kilian Q. Weinberger
摘要
Large language models (LLMs) should undergo rigorous audits to identify potential risks, such as copyright and privacy infringements. Once these risks emerge, timely updates are crucial to remove undesirable responses, ensuring legal and safe model usage. It has spurred recent research into LLM unlearning, focusing on erasing targeted undesirable knowledge without compromising the integrity of other, non-targeted responses. Existing studies have introduced various unlearning objectives to pursue LLM unlearning without necessitating complete retraining. However, each of these objectives has unique properties, and no unified framework is currently available to comprehend them thoroughly. To fill the gap, we propose a toolkit of the gradient effect (G-effect), quantifying the impacts of unlearning objectives on model performance from a gradient perspective. A notable advantage is its broad ability to detail the unlearning impacts from various aspects across instances, updating steps, and LLM layers. Accordingly, the G-effect offers new insights into identifying drawbacks of existing unlearning objectives, further motivating us to explore a series of new solutions for their mitigation and improvements. Finally, we outline promising directions that merit further studies, aiming at contributing to the community to advance this important field. The code is publicly available at: https://github.com/tmlr-group/G-effect .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper22
- LLM Unlearning with LLM BeliefsKemou Li, Qizhou Wang, Yue Wang, Fengpeng Li 等ICLR 2026 · 被引用 20 次
- Self-Destructive Language ModelsYuhui Wang, Rongyi Zhu, Ting WangICLR 2026 · 被引用 14 次
- Constrained Entropic Unlearning: A Primal-Dual Framework for Large Language ModelsTaha Entesari, Arman Hatami, Rinat Khaziev, Anil Ramakrishna 等NeurIPS 2025 · 被引用 12 次
- Explainable LLM Unlearning through ReasoningJunfeng Liao, Qizhou Wang, Shanshan Ye, Xin Yu 等ICLR 2026 · 被引用 8 次
- Attention Smoothing Is All You Need For UnlearningSaleh Zare Zade, Xiangyu Zhou, Sijia Liu, Dongxiao ZhuICLR 2026 · 被引用 7 次
它引用的顶会 Paper11
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- The WMDP Benchmark: Measuring and Reducing Malicious Use with UnlearningNathaniel Li, Alexander Pan, Anjali Gopal, Summer Yue 等ICML 2024 · 被引用 390 次
- Large Language Model UnlearningYuanshun Yao, Xiaojun Xu, Yang LiuNeurIPS 2024 · 被引用 365 次
相关 Paper
- GRU: Mitigating the Trade-off between Unlearning and Retention for LLMsYue Wang, Qizhou Wang, Feng Liu, Wei Huang 等ICML 2025
- Forget to Know, Remember to Use: Context-Aware Unlearning for Large Language ModelsYuefeng Peng, Parnian Afshar, Megan Ganji, Thomas Butler 等ICML 2026 · 被引用 1 次
- Towards Effective Evaluations and Comparisons for LLM Unlearning MethodsQizhou Wang, Bo Han, Puning Yang, Jianing Zhu 等ICLR 2025
- Downgrade to Upgrade: Optimizer Simplification Enhances Robustness in LLM UnlearningYicheng Lang, Yihua Zhang, Chongyu Fan, Changsheng Wang 等ICLR 2026 · 被引用 4 次
- To Each (Textual Sequence) Its Own: Improving Memorized-Data Unlearning in Large Language ModelsGeorge-Octavian Barbulescu, Peter TriantafillouICML 2024 · 被引用 41 次
