A Robust Unlearning Method with Adaptive Knowledge Guidance and Memory Preservation
Jingyuan Tian, Xiaofei Zhou
Abstract
Machine unlearning has emerged as a promising approach to remove specific knowledge from large language models (LLMs), especially for safety-critical applications. However, existing representation-based methods lack guidance for selecting representation locations to unlearn (RMU), thus lacking precision in unlearning, while probability-based methods are vulnerable to fine-tuning attacks which use unrelated and safe data to fine-tune models. To address these problems, this paper presents an adaptive knowledge guidance and memory perturbation mechanisms, called ALMPU (Adaptive Localized Memory Perturbation Unlearning) which addresses the lack of knowledge guidance in representation-based unlearning methods and mitigates the impact of fine-tuning attacks on unlearned models. Specifically, we apply scaling factors to attention heads and select the most sensitive ones as knowledge guidance. Guided by the previous knowledge localization, we integrate enhanced memory perturbation—which forces the model to preserve specific knowledge—into the standard representation-based unlearning process at these sensitive positions. Through this perturbation mechanism, the model achieves more thorough elimination of the target knowledge. By adding interventions to selected attention heads and explicitly optimizing against fine-tuning attacks during the unlearning process, ALMPU creates a controlled divergence from the original model that is inherently resistant to relearning attempts. Experimental evaluation on the WMDP benchmark demonstrates that ALMPU consistently outperforms baseline methods across different scales of fine-tuning attacks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d58a0305-5f8d-4fa5-8610-c5805bbdc2fbBuilds on11
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- The WMDP Benchmark: Measuring and Reducing Malicious Use with UnlearningNathaniel Li, Alexander Pan, Anjali Gopal, Summer Yue et al.ICML 2024 · 390 citations
- Large Language Model UnlearningYuanshun Yao, Xiaojun Xu, Yang LiuNeurIPS 2024 · 365 citations
- LoFiT: Localized Fine-tuning on LLM RepresentationsFangcong Yin, Xi Ye, Greg DurrettNeurIPS 2024 · 74 citations
- On Effects of Steering Latent Representation for Large Language Model UnlearningHuu-Tien Dang, Tin Pham, Hoang Thanh-Tung, Naoya InoueAAAI 2025 · 33 citations
Related papers
- Invariance Makes LLM Unlearning Resilient Even to Unanticipated Downstream Fine-TuningChangsheng Wang, Yihua Zhang, Jinghan Jia, Parikshit Ram et al.ICML 2025
- Elastic Robust Unlearning of Specific Knowledge in Large Language ModelsYize Sui, Jing Ren, Wenjing Yang, Ruochun Jin et al.NeurIPS 2025 · 1 citation
- Model Unlearning via Sparse Autoencoder Subspace Guided ProjectionsXu Wang, Zihao Li, Benyou Wang, Yan Hu et al.EMNLP 2025 · 9 citations
- Unlearning or Obfuscating? Jogging the Memory of Unlearned LLMs via Benign RelearningShengyuan Hu, Yiwei Fu, Steven Z. Wu, Virginia SmithICLR 2025
- Keeping an Eye on LLM Unlearning: The Hidden Risk and RemedyJie Ren, Zhenwei Dai, Xianfeng Tang, Yue Xing et al.NeurIPS 2025 · 11 citations
