Probing Hidden Knowledge Holes in Unlearned LLMs
Myeongseob Ko, Hoang Anh Just, Charles Fleming, Ming Jin, Ruoxi Jia
摘要
Machine unlearning has emerged as a prevalent technical solution for selectively removing unwanted knowledge absorbed during pre-training, without requiring full retraining. While recent unlearning techniques can effectively remove undesirable content without severely compromising performance on standard benchmarks, we find that they may inadvertently create "knowledge holes"-unintended losses of benign knowledge that standard benchmarks fail to capture. To probe where unlearned models reveal knowledge holes, we propose a test case generation framework that explores both immediate neighbors of unlearned content and broader areas of potential failures. Our evaluation demonstrates significant hidden costs of unlearning: up to 98.7% of the test cases yield irrelevant or nonsensical responses from unlearned models, despite being answerable by the pretrained model. These findings necessitate rethinking the conventional approach to evaluating knowledge preservation in unlearning, moving beyond standard, static benchmarks.
- 'Extremely low-quality' refers to answers scoring 1 on a 1-10 scale when evaluated by the LLM-as-a-judge method [Zheng et al., 2023]. These are typically incomplete or gibberish responses. * We use "prompts" and "test cases" interchangeably throughout the paper.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper9
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- TruthfulQA: Measuring How Models Mimic Human FalsehoodsStephanie Lin, Jacob Hilton, Owain EvansACL 2022 · 被引用 3,228 次
- The WMDP Benchmark: Measuring and Reducing Malicious Use with UnlearningNathaniel Li, Alexander Pan, Anjali Gopal, Summer Yue 等ICML 2024 · 被引用 390 次
- Large Language Model UnlearningYuanshun Yao, Xiaojun Xu, Yang LiuNeurIPS 2024 · 被引用 365 次
- Red Teaming Language Models with Language ModelsEthan Perez, Saffron Huang, H. Francis Song, Trevor Cai 等EMNLP 2022 · 被引用 239 次
相关 Paper
- Intrinsic Test of Unlearning Using Parametric Knowledge TracesYihuai Hong, Lei Yu, Haiqin Yang, Shauli Ravfogel 等EMNLP 2025 · 被引用 1 次
- Beyond Superficial Forgetting: Thorough Unlearning Through Knowledge Density Estimation and Block Re-InsertionFeng Guo, Yuntao Wen, Shen Gao, Junshuo Zhang 等AAAI 2026
- Do LLMs Really Forget? Evaluating Unlearning with Knowledge Correlation and Confidence AwarenessRongzhe Wei, Peizhi Niu, Hans Hao-Hsun Hsu, Ruihan Wu 等NeurIPS 2025 · 被引用 20 次
- Catastrophic Failure of LLM Unlearning via QuantizationZhiwei Zhang, Fali Wang, Xiaomin Li, Zongyu Wu 等ICLR 2025
- Unlearning Isn't Invisible: Detecting Unlearning Traces in LLMs from Model OutputsYiwei Chen, Soumyadeep Pal, Yimeng Zhang, Qing Qu 等ICLR 2026 · 被引用 15 次
