Beyond Superficial Forgetting: Thorough Unlearning Through Knowledge Density Estimation and Block Re-Insertion
Feng Guo, Yuntao Wen, Shen Gao, Junshuo Zhang, Shuo Shang
Abstract
Machine unlearning, which selectively removes harmful knowledge from a pre-trained model without retraining from scratch, is crucial for addressing privacy, regulatory compliance, and ethical concerns in Large Language Models (LLMs). However, existing unlearning methods often struggle to thoroughly remove harmful knowledge, leaving residual harmful knowledge that can be easily recovered. To address these limitations, we propose Knowledge Density-Guided Unlearning via Blocks Reinsertion (KUnBR), a novel approach that first identifies layers with rich harmful knowledge and then thoroughly eliminates the harmful knowledge via re-insertion strategy. Our method introduces knowledge density estimation to quantify and locate layers containing the most harmful knowledge, enabling precise unlearning. Additionally, we design a layer re-insertion strategy that extracts and re-inserts harmful knowledge-rich layers into the original LLM, bypassing gradient obstruction caused by cover layers and ensuring effective gradient propagation during unlearning. Extensive experiments conducted on several unlearning and general capability benchmarks demonstrate that KUnBR achieves state-of-the-art forgetting performance while maintaining model utility 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on10
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- TruthfulQA: Measuring How Models Mimic Human FalsehoodsStephanie Lin, Jacob Hilton, Owain EvansACL 2022 · 3,228 citations
- Extracting Training Data from Large Language ModelsNicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski et al.USENIX Security 2021 · 2,866 citations
- The WMDP Benchmark: Measuring and Reducing Malicious Use with UnlearningNathaniel Li, Alexander Pan, Anjali Gopal, Summer Yue et al.ICML 2024 · 390 citations
- Soft Prompt Threats: Attacking Safety Alignment and Unlearning in Open-Source LLMs through the Embedding SpaceLeo Schwinn, David Dobre, Sophie Xhonneux, Gauthier Gidel et al.NeurIPS 2024 · 113 citations
Related papers
- Adaptive Localization of Knowledge Negation for Continual LLM UnlearningAbudukelimu Wuerkaixi, Qizhou Wang, Sen Cui, Wutong Xu et al.ICML 2025
- Unlearning vs. Obfuscation: Are We Truly Removing Knowledge?Guangzhi Sun, Potsawee Manakul, Xiao Zhan, Mark J. F. GalesEMNLP 2025
- LLM-Eraser: Optimizing Large Language Model Unlearning through Selective PruningShengming Zhang, Le Zhang, Jingbo Zhou, Zhi Zheng et al.KDD 2025 · 2 citations
- Large Scale Knowledge WashingYu Wang, Ruihan Wu, Zexue He, Xiusi Chen et al.ICLR 2025
- ZeroUnlearn: Few-Shot Knowledge Unlearning in Large Language ModelsYujie Lin, Chengyi Yang, Zhishang Xiang, YIPING SONG et al.ICML 2026
