LLM-Eraser: Optimizing Large Language Model Unlearning through Selective Pruning
Shengming Zhang, Le Zhang, Jingbo Zhou, Zhi Zheng, Hui Xiong
摘要
We focus on unlearning unwanted knowledge in autoregressive large language models (LLMs) through pruning. Our goal is to selectively remove undesirable information (e.g., harmful responses, privacy-sensitive data) while ensuring the preservation of desirable knowledge (e.g., positive responses and objective facts). Previous approaches use gradient ascent (GA) over undesired knowledge to inversely optimize LLMs, which compromises the model's performance on desired knowledge. To address this limitation, we introduce a novel two-stage approach, named LLM-Eraser, for selectively identifying and editing parameters specifically associated with undesirable knowledge. LLM-Eraser operates in two stages: localization and unlearning. During the localization stage, we utilize neuron scores and trainable soft masks to identify parameters crucial to the undesired knowledge. In the unlearning stage, we prune these identified parameters and apply a selective post-training process to enhance the model's selectiveness. Our experiments, conducted across five task datasets, demonstrate that LLM-Eraser effectively unlearns undesirable knowledge-evidenced by the model's near-random performance on multiple-choice questions related to the erased knowledge-while maintaining high proficiency in desirable knowledge, with an average performance deficit of only 2.5%.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Think and Recall: Layer-Level Prompting for Lifelong Model EditingJinke Wang, Zenan Ying, Qi Liu, Wei Chen 等EMNLP 2025
- DIAA: A Decoding-Efficient Inference Acceleration Approach for On-Device Large Language ModelsHao Tian, Sheng Lu, Fuwen Tian, Guangming Cui 等AAAI 2026
- EPTS: Elastic Post-Training Sparsity for Efficient Large Language Model CompressionKe Xu, Jiaqi Wan, Wenhao Hu, Han Pu 等KDD 2026
它引用的顶会 Paper30
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Supervised Contrastive LearningPrannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna 等NeurIPS 2020 · 被引用 7,049 次
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 被引用 3,415 次
相关 Paper
- Large Scale Knowledge WashingYu Wang, Ruihan Wu, Zexue He, Xiusi Chen 等ICLR 2025
- VL-Eraser: Vacuum Distillation for Machine Unlearning in Vision-Language ModelsYili Wang, Lu Dai, Tairan Huang, Yijie Xu 等CVPR 2026
- Decoding-Unlearning: Fact Forgetting via Entropy-Guided InferenceJingwen Pu, Mingjun Shi, Xinrui Ren, Yizhe Wang 等ACL 2026
- Explainable LLM Unlearning through ReasoningJunfeng Liao, Qizhou Wang, Shanshan Ye, Xin Yu 等ICLR 2026 · 被引用 8 次
- Constrained Entropic Unlearning: A Primal-Dual Framework for Large Language ModelsTaha Entesari, Arman Hatami, Rinat Khaziev, Anil Ramakrishna 等NeurIPS 2025 · 被引用 12 次
