Unlearning vs. Obfuscation: Are We Truly Removing Knowledge?
Guangzhi Sun, Potsawee Manakul, Xiao Zhan, Mark J. F. Gales
摘要
Unlearning has emerged as a critical capability for large language models (LLMs) to support data privacy, regulatory compliance, and ethical AI deployment. Recent techniques often rely on obfuscation by injecting incorrect or irrelevant information to suppress knowledge. Such methods effectively constitute knowledge addition rather than true removal, often leaving models vulnerable to probing. In this paper, we formally distinguish unlearning from obfuscation and introduce a probing-based evaluation framework to assess whether existing approaches genuinely remove targeted information. Moreover, we propose DF-MCQ, a novel unlearning method that flattens the model predictive distribution over automatically generated multiple-choice questions using KLdivergence, effectively removing knowledge about target individuals and triggering appropriate refusal behaviour. Experimental results demonstrate that DF-MCQ achieves unlearning with over 90% refusal rate and a random choice-level uncertainty that is much higher than obfuscation on probing questions. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper13
- Memory-Based Model Editing at ScaleEric Mitchell, Charles Lin, Antoine Bosselut, Christopher D. Manning 等ICML 2022 · 被引用 520 次
- Large Language Model UnlearningYuanshun Yao, Xiaojun Xu, Yang LiuNeurIPS 2024 · 被引用 365 次
- Knowledge Unlearning for Mitigating Privacy Risks in Language ModelsJoel Jang, Dongkeun Yoon, Sohee Yang, Sungmin Cha 等ACL 2023 · 被引用 48 次
- To Each (Textual Sequence) Its Own: Improving Memorized-Data Unlearning in Large Language ModelsGeorge-Octavian Barbulescu, Peter TriantafillouICML 2024 · 被引用 41 次
- Unlearn What You Want to Forget: Efficient Unlearning for LLMsJiaao Chen, Diyi YangEMNLP 2023 · 被引用 40 次
相关 Paper
- Towards Effective Evaluations and Comparisons for LLM Unlearning MethodsQizhou Wang, Bo Han, Puning Yang, Jianing Zhu 等ICLR 2025
- SEPS: A Separability Measure for Robust Unlearning in LLMsWonje Jeung, Sangyeon Yoon, Albert NoEMNLP 2025
- OBLIVIATE: Robust and Practical Machine Unlearning for Large Language ModelsXiaoyu Xu, Minxin Du, Qingqing Ye, Haibo HuEMNLP 2025 · 被引用 1 次
- Beyond Superficial Forgetting: Thorough Unlearning Through Knowledge Density Estimation and Block Re-InsertionFeng Guo, Yuntao Wen, Shen Gao, Junshuo Zhang 等AAAI 2026
- Tool Unlearning for Tool-Augmented LLMsJiali Cheng, Hadi AmiriICML 2025
