Unlearning vs. Obfuscation: Are We Truly Removing Knowledge?
Guangzhi Sun, Potsawee Manakul, Xiao Zhan, Mark J. F. Gales
Abstract
Unlearning has emerged as a critical capability for large language models (LLMs) to support data privacy, regulatory compliance, and ethical AI deployment. Recent techniques often rely on obfuscation by injecting incorrect or irrelevant information to suppress knowledge. Such methods effectively constitute knowledge addition rather than true removal, often leaving models vulnerable to probing. In this paper, we formally distinguish unlearning from obfuscation and introduce a probing-based evaluation framework to assess whether existing approaches genuinely remove targeted information. Moreover, we propose DF-MCQ, a novel unlearning method that flattens the model predictive distribution over automatically generated multiple-choice questions using KLdivergence, effectively removing knowledge about target individuals and triggering appropriate refusal behaviour. Experimental results demonstrate that DF-MCQ achieves unlearning with over 90% refusal rate and a random choice-level uncertainty that is much higher than obfuscation on probing questions. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 71379229-5bd6-4d7d-81bb-2fafcff8b04bCited by top-tier papers1
Ask how each one uses itBuilds on13
- Memory-Based Model Editing at ScaleEric Mitchell, Charles Lin, Antoine Bosselut, Christopher D. Manning et al.ICML 2022 · 520 citations
- Large Language Model UnlearningYuanshun Yao, Xiaojun Xu, Yang LiuNeurIPS 2024 · 365 citations
- Knowledge Unlearning for Mitigating Privacy Risks in Language ModelsJoel Jang, Dongkeun Yoon, Sohee Yang, Sungmin Cha et al.ACL 2023 · 48 citations
- To Each (Textual Sequence) Its Own: Improving Memorized-Data Unlearning in Large Language ModelsGeorge-Octavian Barbulescu, Peter TriantafillouICML 2024 · 41 citations
- Unlearn What You Want to Forget: Efficient Unlearning for LLMsJiaao Chen, Diyi YangEMNLP 2023 · 40 citations
Related papers
- Towards Effective Evaluations and Comparisons for LLM Unlearning MethodsQizhou Wang, Bo Han, Puning Yang, Jianing Zhu et al.ICLR 2025
- SEPS: A Separability Measure for Robust Unlearning in LLMsWonje Jeung, Sangyeon Yoon, Albert NoEMNLP 2025
- OBLIVIATE: Robust and Practical Machine Unlearning for Large Language ModelsXiaoyu Xu, Minxin Du, Qingqing Ye, Haibo HuEMNLP 2025 · 1 citation
- Beyond Superficial Forgetting: Thorough Unlearning Through Knowledge Density Estimation and Block Re-InsertionFeng Guo, Yuntao Wen, Shen Gao, Junshuo Zhang et al.AAAI 2026
- Tool Unlearning for Tool-Augmented LLMsJiali Cheng, Hadi AmiriICML 2025
