On Large Language Model Continual Unlearning
Chongyang Gao, Lixu Wang, Kaize Ding, Chenkai Weng, Xiao Wang, Qi Zhu
摘要
While large language models have demonstrated impressive performance across various domains and tasks, their security issues have become increasingly severe. Machine unlearning has emerged as a representative approach for model safety and security by removing the influence of undesired data on the target model. However, these methods do not sufficiently consider that unlearning requests in real-world scenarios are continuously emerging, especially in the context of LLMs, which may lead to accumulated model utility loss that eventually becomes unacceptable. Moreover, existing LLM unlearning methods often ignore previous data access limitations due to privacy concerns and copyright protection. Without previous data, the utility preservation during unlearning is much harder. To overcome these challenges, we propose the O 3 framework that includes an Orthogonal low-rank adapter (LoRA) for continually unlearning requested data and an Out-Of-Distribution (OOD) detector to measure the similarity between input and unlearning data. The orthogonal LoRA achieves parameter disentanglement among continual unlearning requests. The OOD detector is trained with a novel contrastive entropy loss and utilizes a glocal-aware scoring mechanism. During inference, our O 3 framework can decide whether and to what extent to load the unlearning LoRA based on the OOD detector's predicted similarity between the input and the unlearned knowledge. Notably, O 3 's effectiveness does not rely on any retained data. We conducted extensive experiments on O 3 and state-of-the-art LLM unlearning methods across three tasks and seven datasets. The results indicate that O 3 consistently achieves the best unlearning effectiveness and utility preservation, especially when facing continuous unlearning requests. The source codes can be found at https://github.com/GCYZSL/O3-LLM-UNLEARNING .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- Unlearning Isn't Deletion: Investigating Reversibility of Machine Unlearning in LLMsXiaoyu Xu, Xiang Yue, Yang Liu, Qingqing Ye 等ICML 2026 · 被引用 36 次
- DRAGON: Guard LLM Unlearning in Context via Negative Detection and ReasoningYaxuan Wang, Chris Yuhao Liu, Quan Liu, Jinlong Pang 等ICLR 2026 · 被引用 10 次
- FG-OrIU: Towards Better Forgetting via Feature-Gradient Orthogonality for Incremental UnlearningQian Feng, Jiahang Tu, Mintong Kang, Hanbin Zhao 等ICCV 2025 · 被引用 9 次
- Continual Unlearning for Text-to-Image Diffusion Models: A Regularization PerspectiveJustin Lee, Zheda Mai, Jinsu Yoo, Chongyu Fan 等ICLR 2026 · 被引用 9 次
- Attention Smoothing Is All You Need For UnlearningSaleh Zare Zade, Xiangyu Zhou, Sijia Liu, Dongxiao ZhuICLR 2026 · 被引用 7 次
它引用的顶会 Paper35
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 被引用 3,415 次
- TruthfulQA: Measuring How Models Mimic Human FalsehoodsStephanie Lin, Jacob Hilton, Owain EvansACL 2022 · 被引用 3,228 次
- Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question AnsweringPan Lu, Swaroop Mishra, Tanglin Xia, Liang Qiu 等NeurIPS 2022 · 被引用 2,727 次
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 被引用 2,496 次
相关 Paper
- Adaptive Localization of Knowledge Negation for Continual LLM UnlearningAbudukelimu Wuerkaixi, Qizhou Wang, Sen Cui, Wutong Xu 等ICML 2025
- OBLIVIATE: Robust and Practical Machine Unlearning for Large Language ModelsXiaoyu Xu, Minxin Du, Qingqing Ye, Haibo HuEMNLP 2025 · 被引用 1 次
- Unified Parameter-Efficient Unlearning for LLMsChenlu Ding, Jiancan Wu, Yancheng Yuan, Jinda Lu 等ICLR 2025
- Towards Practical LLM Unlearning: Efficient, Modular, and Retain-FreePeng Liu, Peng-Fei Zhang, Jianfeng Qu, Ximing Li 等WWW 2026
- ALTER: Asymmetric LoRA for Token-Entropy-Guided Unlearning of LLMsXunlei Chen, Jinyu Guo, Yuang Li, Zhaokun Wang 等AAAI 2026 · 被引用 2 次
