Distillation Robustifies Unlearning
Bruce W. Lee, Addie Foote, Alex Infanger, Leni Shor, Harish Kamath, Jacob Goldman-Wetzler, Bryce Woodworth, Alex Cloud, Alexander Matt Turner
摘要
Current LLM unlearning methods are not robust. A few steps of finetuning can revert their effects. We begin by showing that this is true even for an idealized form of unlearning: training to imitate a model that was never trained on unwanted information. This shows that training a model can drastically modify its input-output behavior while leaving its underlying capabilities intact. In light of this dynamic, we show our main result. Training a randomly initialized student on the outputs of an unlearned model transfers behaviors while leaving latent capabilities behind. In short, distillation robustifies unlearning. Based on this result, we propose Unlearn-Noise-Distill-on-Outputs (UNDO), a scalable method that distills an unlearned model into a noised copy of itself. UNDO introduces a tunable tradeoff between compute cost and robustness, establishing a new Pareto frontier on synthetic language and arithmetic tasks. At its strongest setting, UNDO matches the robustness of a model retrained from scratch with perfect data filtering while using only 60-80% of the compute and requiring only 0.01% of the pretraining data to be labeled. We also show that UNDO robustifies unlearning on the more realistic Weapons of Mass Destruction Proxy (WMDP) benchmark. Since distillation is widely used in practice, incorporating an unlearning step beforehand offers a convenient path to robust capability removal.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Deep Ignorance: Filtering Pretraining Data Builds Tamper-Resistant Safeguards into Open-Weight LLMsKyle O'Brien, Stephen Casper, Quentin Anthony, Tomek Korbak 等ICLR 2026 · 被引用 59 次
- On the creation of narrow AI: hierarchy and nonlocality of neural network skillsEric J. Michaud, Asher Parker-Sartori, Max TegmarkNeurIPS 2025 · 被引用 6 次
- Dual-Space Smoothness for Robust and Balanced LLM UnlearningHan Yan, Zheyuan Liu, Meng JiangICLR 2026 · 被引用 5 次
- Modular Pretraining Enables Access ControlEthan Roland, Murat Cubuktepe, Erick Martinez, Stijn Servaes 等ICML 2026
它引用的顶会 Paper41
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Membership Inference Attacks Against Machine Learning ModelsReza Shokri, Marco Stronati, Congzheng Song, Vitaly ShmatikovS&P 2017 · 被引用 5,137 次
- Machine UnlearningLucas Bourtoule, Varun Chandrasekaran, Christopher A. Choquette-Choo, Hengrui Jia 等S&P 2021 · 被引用 1,381 次
- Are Emergent Abilities of Large Language Models a Mirage?Rylan Schaeffer, Brando Miranda, Sanmi KoyejoNeurIPS 2023 · 被引用 796 次
- Certified Data Removal from Machine Learning ModelsChuan Guo, Tom Goldstein, Awni Y. Hannun, Laurens van der MaatenICML 2020 · 被引用 633 次
相关 Paper
- DUET: Distilled LLM Unlearning from an Efficiently Contextualized TeacherYisheng Zhong, Zhengbang Yang, Zhuangdi ZhuICLR 2026 · 被引用 4 次
- Invariance Makes LLM Unlearning Resilient Even to Unanticipated Downstream Fine-TuningChangsheng Wang, Yihua Zhang, Jinghan Jia, Parikshit Ram 等ICML 2025
- Elastic Robust Unlearning of Specific Knowledge in Large Language ModelsYize Sui, Jing Ren, Wenjing Yang, Ruochun Jin 等NeurIPS 2025 · 被引用 1 次
- Reversing the Forget-Retain Objectives: An Efficient LLM Unlearning Framework from Logit DifferenceJiabao Ji, Yujian Liu, Yang Zhang, Gaowen Liu 等NeurIPS 2024 · 被引用 106 次
- Attention Smoothing Is All You Need For UnlearningSaleh Zare Zade, Xiangyu Zhou, Sijia Liu, Dongxiao ZhuICLR 2026 · 被引用 7 次
