Keeping an Eye on LLM Unlearning: The Hidden Risk and Remedy
Jie Ren, Zhenwei Dai, Xianfeng Tang, Yue Xing, Shenglai Zeng, Jingying Zeng, Qiankun Peng, Samarth Varshney, Suhang Wang, Qi He, Charu Aggarwal, Hui Liu
Abstract
Although Large Language Models (LLMs) have demonstrated impressive capabilities across a wide range of tasks, growing concerns have emerged over the misuse of sensitive, copyrighted, or harmful data during training. To address these concerns, unlearning techniques have been developed to remove the influence of specific data without retraining from scratch. However, this paper reveals a critical vulnerability in fine-tuning-based unlearning: a malicious user can craft a manipulated forgetting request that stealthily degrades the model's utility for benign users. We demonstrate this risk through a red-teaming Stealthy Attack (SA), which is inspired by two key limitations of existing unlearning-the inability to constrain the scope of unlearning effect and the failure to distinguish benign tokens from unlearning signals. Prior work has shown that unlearned models tend to memorize forgetting data as unlearning signals, and respond with hallucinations or feigned ignorance when unlearning signals appear in the input. By subtly increasing the presence of common benign tokens in the forgetting data, SA enhances the connection between benign tokens and unlearning signals. As a result, when normal users include such tokens in their prompts, the model exhibits unlearning behaviors, leading to unintended utility degradation. To address this vulnerability, we propose Scope-aware Unlearning (SU), a lightweight enhancement that introduces a scope term into the unlearning objective, encouraging the model to localize the forgetting effect. Our method requires no additional data processing, integrates seamlessly with existing fine-tuning frameworks, and significantly improves robustness against SA. Extensive experiments validate the effectiveness of both SA and SU. Our code is available at github.com/renjie3/sa_su.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f3930c1e-e644-48ef-af28-dabaec7abc5eCited by top-tier papers2
- LLM Unlearning with LLM BeliefsKemou Li, Qizhou Wang, Yue Wang, Fengpeng Li et al.ICLR 2026 · 20 citations
- Hair-Trigger Alignment: Black-Box Evaluation Cannot Guarantee Post-Update AlignmentYavuz Faruk Bakman, Duygu Nur Yaldiz, Eleni Triantafillou, Peter Kairouz et al.ICML 2026 · 2 citations
Builds on18
- Machine UnlearningLucas Bourtoule, Varun Chandrasekaran, Christopher A. Choquette-Choo, Hengrui Jia et al.S&P 2021 · 1,381 citations
- The WMDP Benchmark: Measuring and Reducing Malicious Use with UnlearningNathaniel Li, Alexander Pan, Anjali Gopal, Summer Yue et al.ICML 2024 · 390 citations
- Large Language Model UnlearningYuanshun Yao, Xiaojun Xu, Yang LiuNeurIPS 2024 · 365 citations
- Towards Unbounded Machine UnlearningMeghdad Kurmanji, Peter Triantafillou, Jamie Hayes, Eleni TriantafillouNeurIPS 2023 · 363 citations
- Simplicity Prevails: Rethinking Negative Preference Optimization for LLM UnlearningChongyu Fan, Jiancheng Liu, Licong Lin, Jinghan Jia et al.NeurIPS 2025 · 182 citations
Related papers
- SUA: Stealthy Multimodal Large Language Model Unlearning AttackXianren Zhang, Hui Liu, Delvin Ce Zhang, Xianfeng Tang et al.EMNLP 2025
- OBLIVIATE: Robust and Practical Machine Unlearning for Large Language ModelsXiaoyu Xu, Minxin Du, Qingqing Ye, Haibo HuEMNLP 2025 · 1 citation
- Forget to Flourish: Leveraging Machine-Unlearning on Pretrained Language Models for Privacy LeakageMd. Rafi Ur Rashid, Jing Liu, Toshiaki Koike-Akino, Ye Wang et al.AAAI 2025 · 17 citations
- Refusal Is Not an Option: Unlearning Safety Alignment of Large Language ModelsMinkyoo Song, Hanna Kim, Jaehan Kim, Seungwon Shin et al.USENIX Security 2025
- Attention Smoothing Is All You Need For UnlearningSaleh Zare Zade, Xiangyu Zhou, Sijia Liu, Dongxiao ZhuICLR 2026 · 7 citations
