LLM Unlearning via Neural Activation Redirection
William F. Shen, Xinchi Qiu, Meghdad Kurmanji, Alex Iacob, Lorenzo Sani, Yihong Chen, Nicola Cancedda, Nicholas D. Lane
Abstract
The ability to selectively remove knowledge from LLMs is highly desirable. However, existing methods often struggle with balancing unlearning efficacy and retain model utility, and lack controllability at inference time to emulate base model behavior as if it had never seen the unlearned data. In this paper, we propose LUNAR, a novel unlearning method grounded in the Linear Representation Hypothesis and operates by redirecting the representations of unlearned data to activation regions that expresses its inability to answer. We show that contrastive features are not a prerequisite for effective activation redirection, and LUNAR achieves state-of-the-art unlearning performance and superior controllability. Specifically, LUNAR achieves between 2.9× and 11.7× improvement in the combined unlearning efficacy and model utility score (Deviation Score) across various base models and generates coherent, contextually appropriate responses post-unlearning. Moreover, LUNAR effectively reduces parameter updates to a single down-projection matrix, a novel design that significantly enhances efficiency by 20× and robustness. Finally, we demonstrate that LUNAR is robust to white-box adversarial attacks and versatile in real-world scenarios, including handling sequential unlearning requests.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f53b5b49-b964-4e5b-8ca0-9fb5cfbe9a7aCited by top-tier papers11
- LLM Unlearning with LLM BeliefsKemou Li, Qizhou Wang, Yue Wang, Fengpeng Li et al.ICLR 2026 · 20 citations
- DRAGON: Guard LLM Unlearning in Context via Negative Detection and ReasoningYaxuan Wang, Chris Yuhao Liu, Quan Liu, Jinlong Pang et al.ICLR 2026 · 10 citations
- Divergence Decoding: Inference-Time Unlearning via Auxiliary ModelsHumzah Merchant, Bradford LevyICML 2026 · 2 citations
- Distinguishable Deletion: Unifying Knowledge Erasure and Refusal for Large Language Model UnlearningPuning Yang, Junchi Yu, Qizhou Wang, Phil Torr et al.ICML 2026 · 1 citation
- Less is More: Geometric Unlearning for LLMs with Minimal Data DisclosureChenchen Tan, Xinghao Li, Shujie Cui, Youyang Qu et al.ICML 2026
Builds on24
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 3,415 citations
- Refusal in Language Models Is Mediated by a Single DirectionAndy Arditi, Oscar Obeso, Aaquib Syed, Daniel Paleka et al.NeurIPS 2024 · 1,166 citations
- Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!Xiangyu Qi, Yi Zeng, Tinghao Xie, Pin-Yu Chen et al.ICLR 2024 · 1,104 citations
- AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language ModelsXiaogeng Liu, Nan Xu, Muhao Chen, Chaowei XiaoICLR 2024 · 722 citations
Related papers
- DUET: Distilled LLM Unlearning from an Efficiently Contextualized TeacherYisheng Zhong, Zhengbang Yang, Zhuangdi ZhuICLR 2026 · 4 citations
- FALCON: Fine-grained Activation Manipulation by Contrastive Orthogonal Unalignment for Large Language ModelJinwei Hu, Zhenglin Huang, Xiangyu Yin, Wenjie Ruan et al.NeurIPS 2025 · 3 citations
- On Effects of Steering Latent Representation for Large Language Model UnlearningHuu-Tien Dang, Tin Pham, Hoang Thanh-Tung, Naoya InoueAAAI 2025 · 33 citations
- Elastic Robust Unlearning of Specific Knowledge in Large Language ModelsYize Sui, Jing Ren, Wenjing Yang, Ruochun Jin et al.NeurIPS 2025 · 1 citation
- CoUn: Empowering Machine Unlearning via Contrastive LearningYasser H. Khalil, Mehdi Setayesh, Hongliang LiNeurIPS 2025 · 4 citations
