LLM Unlearning via Neural Activation Redirection
William F. Shen, Xinchi Qiu, Meghdad Kurmanji, Alex Iacob, Lorenzo Sani, Yihong Chen, Nicola Cancedda, Nicholas D. Lane
摘要
The ability to selectively remove knowledge from LLMs is highly desirable. However, existing methods often struggle with balancing unlearning efficacy and retain model utility, and lack controllability at inference time to emulate base model behavior as if it had never seen the unlearned data. In this paper, we propose LUNAR, a novel unlearning method grounded in the Linear Representation Hypothesis and operates by redirecting the representations of unlearned data to activation regions that expresses its inability to answer. We show that contrastive features are not a prerequisite for effective activation redirection, and LUNAR achieves state-of-the-art unlearning performance and superior controllability. Specifically, LUNAR achieves between 2.9× and 11.7× improvement in the combined unlearning efficacy and model utility score (Deviation Score) across various base models and generates coherent, contextually appropriate responses post-unlearning. Moreover, LUNAR effectively reduces parameter updates to a single down-projection matrix, a novel design that significantly enhances efficiency by 20× and robustness. Finally, we demonstrate that LUNAR is robust to white-box adversarial attacks and versatile in real-world scenarios, including handling sequential unlearning requests.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- LLM Unlearning with LLM BeliefsKemou Li, Qizhou Wang, Yue Wang, Fengpeng Li 等ICLR 2026 · 被引用 20 次
- DRAGON: Guard LLM Unlearning in Context via Negative Detection and ReasoningYaxuan Wang, Chris Yuhao Liu, Quan Liu, Jinlong Pang 等ICLR 2026 · 被引用 10 次
- Divergence Decoding: Inference-Time Unlearning via Auxiliary ModelsHumzah Merchant, Bradford LevyICML 2026 · 被引用 2 次
- Distinguishable Deletion: Unifying Knowledge Erasure and Refusal for Large Language Model UnlearningPuning Yang, Junchi Yu, Qizhou Wang, Phil Torr 等ICML 2026 · 被引用 1 次
- Less is More: Geometric Unlearning for LLMs with Minimal Data DisclosureChenchen Tan, Xinghao Li, Shujie Cui, Youyang Qu 等ICML 2026
它引用的顶会 Paper24
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 被引用 3,415 次
- Refusal in Language Models Is Mediated by a Single DirectionAndy Arditi, Oscar Obeso, Aaquib Syed, Daniel Paleka 等NeurIPS 2024 · 被引用 1,166 次
- Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!Xiangyu Qi, Yi Zeng, Tinghao Xie, Pin-Yu Chen 等ICLR 2024 · 被引用 1,104 次
- AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language ModelsXiaogeng Liu, Nan Xu, Muhao Chen, Chaowei XiaoICLR 2024 · 被引用 722 次
相关 Paper
- DUET: Distilled LLM Unlearning from an Efficiently Contextualized TeacherYisheng Zhong, Zhengbang Yang, Zhuangdi ZhuICLR 2026 · 被引用 4 次
- FALCON: Fine-grained Activation Manipulation by Contrastive Orthogonal Unalignment for Large Language ModelJinwei Hu, Zhenglin Huang, Xiangyu Yin, Wenjie Ruan 等NeurIPS 2025 · 被引用 3 次
- On Effects of Steering Latent Representation for Large Language Model UnlearningHuu-Tien Dang, Tin Pham, Hoang Thanh-Tung, Naoya InoueAAAI 2025 · 被引用 33 次
- Elastic Robust Unlearning of Specific Knowledge in Large Language ModelsYize Sui, Jing Ren, Wenjing Yang, Ruochun Jin 等NeurIPS 2025 · 被引用 1 次
- CoUn: Empowering Machine Unlearning via Contrastive LearningYasser H. Khalil, Mehdi Setayesh, Hongliang LiNeurIPS 2025 · 被引用 4 次
