Leak@: Unlearning Does Not Make LLMs Forget Under Probabilistic Decoding
Hadi Reisizadeh, Jiajun Ruan, Yiwei Chen, Soumyadeep Pal, Sijia Liu, Mingyi Hong
Abstract
Unlearning in large language models (LLMs) is critical for regulatory compliance and for building ethical generative AI systems that avoid producing private, toxic, illegal, or copyrighted content. Despite rapid progress, in this work, we show that almost all existing unlearning methods fail to achieve true forgetting in practice. Specifically, while evaluations of these unlearned' models under deterministic (greedy) decoding often suggest successful knowledge removal using standard benchmarks, we show that sensitive information reliably resurfaces when models are sampled with standard probabilistic decoding. To rigorously capture this vulnerability, we introduce leak@, a new meta-evaluation metric that quantifies the likelihood of forgotten knowledge reappearing when generating $k$ samples from the model under realistic decoding strategies. Using three widely adopted benchmarks, TOFU, MUSE, and WMDP, we conduct the first large-scale, systematic study of unlearning reliability using leak@ metric. Our findings demonstrate that knowledge leakage persists across methods and tasks, underscoring that current state-of-the-art (SOTA) unlearning techniques provide only limited forgetting. We propose an algorithm, termed Robust Unlearning under LEak@$k$ metric (RULE) to address this concern. We demonstrate that RULEprovides an unlearned model for TOFU benchmark with no information leakage for a large number of generation samples. On the MUSE benchmark,RULEoutperforms SOTA unlearning methods under theleak@` metric across most sampling budgets . Codes are available at https://github.com/OptimAI-Lab/Leak-k.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 98559329-2e12-4861-b674-45e6fbfa8f2fCited by top-tier papers2
- Divergence Decoding: Inference-Time Unlearning via Auxiliary ModelsHumzah Merchant, Bradford LevyICML 2026 · 2 citations
- Distinguishable Deletion: Unifying Knowledge Erasure and Refusal for Large Language Model UnlearningPuning Yang, Junchi Yu, Qizhou Wang, Phil Torr et al.ICML 2026 · 1 citation
Builds on17
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- The WMDP Benchmark: Measuring and Reducing Malicious Use with UnlearningNathaniel Li, Alexander Pan, Anjali Gopal, Summer Yue et al.ICML 2024 · 390 citations
- Large Language Model UnlearningYuanshun Yao, Xiaojun Xu, Yang LiuNeurIPS 2024 · 365 citations
Related papers
- MUSE: Machine Unlearning Six-Way Evaluation for Language ModelsWeijia Shi, Jaechan Lee, Yangsibo Huang, Sadhika Malladi et al.ICLR 2025
- Constrained Entropic Unlearning: A Primal-Dual Framework for Large Language ModelsTaha Entesari, Arman Hatami, Rinat Khaziev, Anil Ramakrishna et al.NeurIPS 2025 · 12 citations
- R-TOFU: Unlearning in Large Reasoning ModelsSangyeon Yoon, Wonje Jeung, Albert NoEMNLP 2025 · 1 citation
- Large Scale Knowledge WashingYu Wang, Ruihan Wu, Zexue He, Xiusi Chen et al.ICLR 2025
- Multilingual Unlearning in LLMs: Transfer, Dynamics, and ReversibilityChaoyi Xiang, Olga Ohrimenko, Benjamin Rubinstein, Lea FrermannICML 2026
