LLM Unlearning with LLM Beliefs
Kemou Li, Qizhou Wang, Yue Wang, Fengpeng Li, Jun Liu, Bo Han, Jiantao Zhou
Abstract
Large language models trained on vast corpora inherently risk memorizing sensitive or harmful content, which may later resurface in their outputs. Prevailing unlearning methods generally rely on gradient ascent and its variants to lower the probability of specific target responses. However, we find that this strategy induces a critical side effect: probability mass is redistributed into high-likelihood regions, often corresponding to semantically related rephrasings of the targets. We refer to this as the squeezing effect, which explains why many methods yield merely spurious unlearning, a problem further obscured by automated metrics (e.g., ROUGE, truth ratio) that misreport actual success. To address this, we propose a bootstrapping (BS) framework that explicitly links the squeezing effect with the model's own high-confidence generations, namely its model beliefs. Since model beliefs inherently capture the very high-likelihood regions where probability mass is squeezed, incorporating them into the unlearning objective directly counters the squeezing effect. By jointly suppressing both target responses and model beliefs, BS-T (token) attenuates high-probability tokens, whereas BS-S (sequence) removes entire high-confidence generations, together achieving more thorough forgetting while preserving utility. Extensive experiments on diverse benchmarks confirm the effectiveness of our approach, with code merged to OpenUnlearning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 94a2b24e-1e88-4e55-90d5-ead3b8eccc8fCited by top-tier papers8
- Co-rewarding: Stable Self-supervised RL for Eliciting Reasoning in Large Language ModelsZizhuo Zhang, Jianing Zhu, Xinmu Ge, Zihua Zhao et al.ICLR 2026 · 16 citations
- Explainable LLM Unlearning through ReasoningJunfeng Liao, Qizhou Wang, Shanshan Ye, Xin Yu et al.ICLR 2026 · 8 citations
- AEGIS: Adversarial Target-Guided Retention-Data-Free Robust Concept Erasure from Diffusion ModelsFengpeng Li, Kemou Li, Qizhou Wang, Bo Han et al.ICLR 2026 · 6 citations
- Towards Understanding Valuable Preference Data for Large Language Model AlignmentZizhuo Zhang, Qizhou Wang, Shanshan Ye, Jianing Zhu et al.ICLR 2026 · 6 citations
- Distinguishable Deletion: Unifying Knowledge Erasure and Refusal for Large Language Model UnlearningPuning Yang, Junchi Yu, Qizhou Wang, Phil Torr et al.ICML 2026 · 1 citation
Builds on47
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Be Your Own Teacher: Improve the Performance of Convolutional Neural Networks via Self DistillationLinfeng Zhang, Jiebo Song, Anni Gao, Jingwei Chen et al.ICCV 2019 · 1,069 citations
- Certified Data Removal from Machine Learning ModelsChuan Guo, Tom Goldstein, Awni Y. Hannun, Laurens van der MaatenICML 2020 · 633 citations
- Remember What You Want to Forget: Algorithms for Machine UnlearningAyush Sekhari, Jayadev Acharya, Gautam Kamath, Ananda Theertha SureshNeurIPS 2021 · 516 citations
Related papers
- Attention Smoothing Is All You Need For UnlearningSaleh Zare Zade, Xiangyu Zhou, Sijia Liu, Dongxiao ZhuICLR 2026 · 7 citations
- To Each (Textual Sequence) Its Own: Improving Memorized-Data Unlearning in Large Language ModelsGeorge-Octavian Barbulescu, Peter TriantafillouICML 2024 · 41 citations
- Forget to Know, Remember to Use: Context-Aware Unlearning for Large Language ModelsYuefeng Peng, Parnian Afshar, Megan Ganji, Thomas Butler et al.ICML 2026 · 1 citation
- Selective Forgetting: Advancing Machine Unlearning Techniques and Evaluation in Language ModelsLingzhi Wang, Xingshan Zeng, Jinsong Guo, Kam-Fai Wong et al.AAAI 2025 · 43 citations
- A Closer Look at Machine Unlearning for Large Language ModelsXiaojian Yuan, Tianyu Pang, Chao Du, Kejiang Chen et al.ICLR 2025
