Forget to Know, Remember to Use: Context-Aware Unlearning for Large Language Models
Yuefeng Peng, Parnian Afshar, Megan Ganji, Thomas Butler, Amir Houmansadr, Mingxian Wang, Dezhi Hong
Abstract
Large language models can memorize information that must be removed-ranging from copyright-sensitive content (e.g., book chapters) to personally identifiable information (e.g., income)-to ensure responsible and compliant behavior. Unlearning has emerged as an efficient alternative to full retraining, aiming to remove specific knowledge. However, users may still expect model to leverage the removed information when it is re-introduced in the prompt. Existing evaluations of unlearning methods focus on (1) the extent of forgetting of the target knowledge (forget set) and ( 2 ) performance preservation on the retain set (i.e., utility), but overlook this critical usability dimension. Through a systematic evaluation of six state-of-the-art unlearning methods, we show that they consistently degrade such contextual utility-the model's ability to use forgotten knowledge when it is provided in context. To address this, we augment unlearning objectives with a plug-in term that explicitly preserves contextual utility. Extensive experiments demonstrate that our approach restores contextual utility to near original levels while still maintaining effective forgetting and retain-set utility.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 14edeaa2-7ca1-4dba-95ff-edfa516b2fc4Cited by top-tier papers1
Ask how each one uses itBuilds on19
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Extracting Training Data from Large Language ModelsNicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski et al.USENIX Security 2021 · 2,866 citations
Related papers
- OFMU: Optimization-Driven Framework for Machine UnlearningSadia Asif, Mohammad Mohammadi AmiriICLR 2026 · 4 citations
- Catastrophic Failure of LLM Unlearning via QuantizationZhiwei Zhang, Fali Wang, Xiaomin Li, Zongyu Wu et al.ICLR 2025
- Reversing the Forget-Retain Objectives: An Efficient LLM Unlearning Framework from Logit DifferenceJiabao Ji, Yujian Liu, Yang Zhang, Gaowen Liu et al.NeurIPS 2024 · 106 citations
- Reinforcement UnlearningDayong Ye, Tianqing Zhu, Congcong Zhu, Derui Wang et al.NDSS 2025
- RULE: Reinforcement UnLEarning Achieves Forget-retain Pareto OptimalityChenlong Zhang, Zhuoran Jin, Hongbang Yuan, Jiaheng Wei et al.NeurIPS 2025 · 15 citations
