A Fully Probabilistic Perspective on Large Language Model Unlearning: Evaluation and Optimization
Anda Cheng, Wei Huang, Yinggui Wang
Abstract
Large Language Model Unlearning (LLMU) is a promising way to remove private or sensitive information from large language models. However, the comprehensive evaluation of LLMU remains underexplored. The dominant deterministic evaluation can yield overly optimistic assessments of unlearning efficacy. To mitigate this, we propose a Fully Probabilistic Evaluation (FPE) framework that incorporates input and output distributions in LLMU evaluation. FPE obtains a probabilistic evaluation result by querying unlearned models with various semantically similar inputs and multiple sampling attempts. We introduce an Input Distribution Sampling method in FPE to select high-quality inputs, enabling a stricter measure of information leakage risks. Furthermore, we introduce a Contrastive Embedding Loss (CEL) to advance the performance of LLMU. CEL employs contrastive learning to distance latent representations of unlearned samples from adaptively clustered contrast samples while aligning them with random vectors, leading to improved efficacy and robustness for LLMU. Our experiments show that FPE uncovers more unlearned information leakage risks than prior evaluation methods, and CEL improves unlearning effectiveness by at least 50.1% and robustness by at least 37.2% on Llama-2-7B while retaining high model utility.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cac65422-34b1-4aeb-9249-a3aa1ba0e08fCited by top-tier papers1
Ask how each one uses itBuilds on10
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Large Language Model UnlearningYuanshun Yao, Xiaojun Xu, Yang LiuNeurIPS 2024 · 365 citations
- Can Sensitive Information Be Deleted From LLMs? Objectives for Defending Against Extraction AttacksVaidehi Patil, Peter Hase, Mohit BansalICLR 2024 · 167 citations
- Soft Prompt Threats: Attacking Safety Alignment and Unlearning in Open-Source LLMs through the Embedding SpaceLeo Schwinn, David Dobre, Sophie Xhonneux, Gauthier Gidel et al.NeurIPS 2024 · 113 citations
- Knowledge Unlearning for Mitigating Privacy Risks in Language ModelsJoel Jang, Dongkeun Yoon, Sohee Yang, Sungmin Cha et al.ACL 2023 · 48 citations
Related papers
- SUA: Stealthy Multimodal Large Language Model Unlearning AttackXianren Zhang, Hui Liu, Delvin Ce Zhang, Xianfeng Tang et al.EMNLP 2025
- SALMUBench: A Benchmark for Sensitive Association-Level Multimodal UnlearningCai Selvas-Sala, Lei Kang, Lluís GómezCVPR 2026 · 4 citations
- Constrained Entropic Unlearning: A Primal-Dual Framework for Large Language ModelsTaha Entesari, Arman Hatami, Rinat Khaziev, Anil Ramakrishna et al.NeurIPS 2025 · 12 citations
- A Probabilistic Perspective on Unlearning and Alignment for Large Language ModelsYan Scholten, Stephan Günnemann, Leo SchwinnICLR 2025
- Leak@: Unlearning Does Not Make LLMs Forget Under Probabilistic DecodingHadi Reisizadeh, Jiajun Ruan, Yiwei Chen, Soumyadeep Pal et al.ICML 2026
