Revisiting Who's Harry Potter: Towards Targeted Unlearning from a Causal Intervention Perspective
Yujian Liu, Yang Zhang, Tommi S. Jaakkola, Shiyu Chang
摘要
This paper investigates Who's Harry Potter (WHP), a pioneering yet insufficiently understood method for LLM unlearning. We explore it in two steps. First, we introduce a new task of LLM targeted unlearning, where given an unlearning target (e.g., a person) and some unlearning documents, we aim to unlearn only the information about the target, rather than everything in the unlearning documents. We further argue that a successful unlearning should satisfy criteria such as not outputting gibberish, not fabricating facts about the unlearning target, and not releasing factual information under jailbreak attacks. Second, we construct a causal intervention framework for targeted unlearning, where the knowledge of the unlearning target is modeled as a confounder between LLM input and output, and the unlearning process as a deconfounding process. This framework justifies and extends WHP, deriving a simple unlearning algorithm that includes WHP as a special case. Experiments on existing and new datasets show that our approach, without explicitly optimizing for the aforementioned criteria, achieves competitive performance in all of them. Our code is available at https://github.com/ UCSB-NLP-Chang/causal_unlearn.git . Q: Where was Wilhelm Wattenbach born? A: I don't have his personal information.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Learning-Time Encoding Shapes Unlearning in LLMsRuihan Wu, Konstantin Garov, Kamalika ChaudhuriICLR 2026 · 被引用 2 次
- ALTER: Asymmetric LoRA for Token-Entropy-Guided Unlearning of LLMsXunlei Chen, Jinyu Guo, Yuang Li, Zhaokun Wang 等AAAI 2026 · 被引用 2 次
- OBLIVIATE: Robust and Practical Machine Unlearning for Large Language ModelsXiaoyu Xu, Minxin Du, Qingqing Ye, Haibo HuEMNLP 2025 · 被引用 1 次
- Privacy and Safety Experiences and Concerns of US Women Using Generative AI for Seeking Sexual and Reproductive Health InformationIna Kaleva, Xiao Zhan, Ruba Abu-Salma, Jose SuchCHI 2026 · 被引用 1 次
- CiPO: Counterfactual Unlearning for Large Reasoning Models through Iterative Preference OptimizationJunyi Li, Yongqiang Chen, Ningning DingACL 2026 · 被引用 1 次
它引用的顶会 Paper28
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Neural Text Generation With Unlikelihood TrainingSean Welleck, Ilia Kulikov, Stephen Roller, Emily Dinan 等ICLR 2020 · 被引用 683 次
- Certified Data Removal from Machine Learning ModelsChuan Guo, Tom Goldstein, Awni Y. Hannun, Laurens van der MaatenICML 2020 · 被引用 633 次
- Erasing Concepts from Diffusion ModelsRohit Gandikota, Joanna Materzynska, Jaden Fiotto-Kaufman, David BauICCV 2023 · 被引用 536 次
- Large Language Model UnlearningYuanshun Yao, Xiaojun Xu, Yang LiuNeurIPS 2024 · 被引用 365 次
相关 Paper
- Unlearning or Obfuscating? Jogging the Memory of Unlearned LLMs via Benign RelearningShengyuan Hu, Yiwei Fu, Steven Z. Wu, Virginia SmithICLR 2025
- Explainable LLM Unlearning through ReasoningJunfeng Liao, Qizhou Wang, Shanshan Ye, Xin Yu 等ICLR 2026 · 被引用 8 次
- MUSE: Machine Unlearning Six-Way Evaluation for Language ModelsWeijia Shi, Jaechan Lee, Yangsibo Huang, Sadhika Malladi 等ICLR 2025
- Large Scale Knowledge WashingYu Wang, Ruihan Wu, Zexue He, Xiusi Chen 等ICLR 2025
- Unlearning vs. Obfuscation: Are We Truly Removing Knowledge?Guangzhi Sun, Potsawee Manakul, Xiao Zhan, Mark J. F. GalesEMNLP 2025
