WAGLE: Strategic Weight Attribution for Effective and Modular Unlearning in Large Language Models
Jinghan Jia, Jiancheng Liu, Yihua Zhang, Parikshit Ram, Nathalie Baracaldo, Sijia Liu
Abstract
The need for effective unlearning mechanisms in large language models (LLMs) is increasingly urgent, driven by the necessity to adhere to data regulations and foster ethical generative AI practices. Despite growing interest of LLM unlearning, much of the existing research has focused on varied unlearning method designs to boost effectiveness and efficiency. However, the inherent relationship between model weights and LLM unlearning has not been extensively examined. In this paper, we systematically explore how model weights interact with unlearning processes in LLMs and we design the weight attribution-guided LLM unlearning method, WAGLE, which unveils the interconnections between 'influence' of weights and 'influence' of data to forget and retain in LLM generation. By strategically guiding the LLM unlearning across different types of unlearning methods and tasks, WAGLE can erase the undesired content, while maintaining the performance of the original tasks. We refer to the weight attribution-guided LLM unlearning method as WAGLE, which unveils the interconnections between 'influence' of weights and 'influence' of data to forget and retain in LLM generation. Our extensive experiments show that WAGLE boosts unlearning performance across a range of LLM unlearning methods such as gradient difference and (negative) preference optimization, applications such as fictitious unlearning, malicious use prevention, and copyrighted information removal, and models including Zephyr-7b-beta and Llama2-7b. To the best of our knowledge, our work offers the first principled method for attributing and pinpointing the influential weights in enhancing LLM unlearning. It stands in contrast to previous methods that lack weight attribution and simpler weight attribution techniques.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fbf72940-3049-4704-818d-67536dd14101Cited by top-tier papers11
- Unlearning Isn't Deletion: Investigating Reversibility of Machine Unlearning in LLMsXiaoyu Xu, Xiang Yue, Yang Liu, Qingqing Ye et al.ICML 2026 · 36 citations
- Unlearning Isn't Invisible: Detecting Unlearning Traces in LLMs from Model OutputsYiwei Chen, Soumyadeep Pal, Yimeng Zhang, Qing Qu et al.ICLR 2026 · 15 citations
- Downgrade to Upgrade: Optimizer Simplification Enhances Robustness in LLM UnlearningYicheng Lang, Yihua Zhang, Chongyu Fan, Changsheng Wang et al.ICLR 2026 · 4 citations
- FALCON: Fine-grained Activation Manipulation by Contrastive Orthogonal Unalignment for Large Language ModelJinwei Hu, Zhenglin Huang, Xiangyu Yin, Wenjie Ruan et al.NeurIPS 2025 · 3 citations
- Randomized Antipodal Search Done Right for Data Pareto Improvement of LLM UnlearningZiwen Liu, Huawei Lin, Yide Ran, Denghui Zhang et al.ICLR 2026 · 2 citations
Builds on41
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 3,415 citations
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 3,037 citations
- Machine UnlearningLucas Bourtoule, Varun Chandrasekaran, Christopher A. Choquette-Choo, Hengrui Jia et al.S&P 2021 · 1,381 citations
Related papers
- De-attribute to Forget for LLM UnlearningXinyang Lu, Jiabao Pan, Rachael Hwee Ling Sim, See-Kiong Ng et al.ICML 2026
- SOUL: Unlocking the Power of Second-Order Optimization for LLM UnlearningJinghan Jia, Yihua Zhang, Yimeng Zhang, Jiancheng Liu et al.EMNLP 2024 · 13 citations
- Simplicity Prevails: Rethinking Negative Preference Optimization for LLM UnlearningChongyu Fan, Jiancheng Liu, Licong Lin, Jinghan Jia et al.NeurIPS 2025 · 182 citations
- Explainable LLM Unlearning through ReasoningJunfeng Liao, Qizhou Wang, Shanshan Ye, Xin Yu et al.ICLR 2026 · 8 citations
- GRU: Mitigating the Trade-off between Unlearning and Retention for LLMsYue Wang, Qizhou Wang, Feng Liu, Wei Huang et al.ICML 2025
