Erasing Conceptual Knowledge from Language Models
Rohit Gandikota, Sheridan Feucht, Samuel Marks, David Bau
摘要
In this work, we introduce Erasure of Language Memory (ELM), a principled approach to concept-level unlearning that operates by matching distributions defined by the model's own introspective classification capabilities. Our key insight is that effective unlearning should leverage the model's ability to evaluate its own knowledge, using the language model itself as a classifier to identify and reduce the likelihood of generating content related to undesired concepts. ELM applies this framework to create targeted low-rank updates that reduce generation probabilities for concept-specific content while preserving the model's broader capabilities. We demonstrate ELM's efficacy on biosecurity, cybersecurity, and literary domain erasure tasks. Comparative evaluation reveals that ELM-modified models achieve near-random performance on assessments targeting erased concepts, while simultaneously preserving generation coherence, maintaining benchmark performance on unrelated tasks, and exhibiting strong robustness to adversarial attacks. Our code, data, and trained models are available at https://elm.baulab.info
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- Deep Ignorance: Filtering Pretraining Data Builds Tamper-Resistant Safeguards into Open-Weight LLMsKyle O'Brien, Stephen Casper, Quentin Anthony, Tomek Korbak 等ICLR 2026 · 被引用 59 次
- Measuring Chain of Thought Faithfulness by Unlearning Reasoning StepsMartin Tutek, Fateme Hashemi Chaleshtori, Ana Marasovic, Yonatan BelinkovEMNLP 2025 · 被引用 37 次
- Unlearning Isn't Deletion: Investigating Reversibility of Machine Unlearning in LLMsXiaoyu Xu, Xiang Yue, Yang Liu, Qingqing Ye 等ICML 2026 · 被引用 36 次
- Obfuscated Activations Bypass LLM Latent-Space DefensesLuke Bailey, Alex Serrano, Abhay Sheshadri, Mikhail Seleznyov 等ICLR 2026 · 被引用 28 次
- Hubble: a Model Suite to Advance the Study of LLM MemorizationJohnny Wei, Ameya Godbole, Mohammad Aflah Khan, Ryan Yixiang Wang 等ICLR 2026 · 被引用 22 次
它引用的顶会 Paper19
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 被引用 3,415 次
- Erasing Concepts from Diffusion ModelsRohit Gandikota, Joanna Materzynska, Jaden Fiotto-Kaufman, David BauICCV 2023 · 被引用 536 次
- The WMDP Benchmark: Measuring and Reducing Malicious Use with UnlearningNathaniel Li, Alexander Pan, Anjali Gopal, Summer Yue 等ICML 2024 · 被引用 390 次
- Large Language Model UnlearningYuanshun Yao, Xiaojun Xu, Yang LiuNeurIPS 2024 · 被引用 365 次
相关 Paper
- Intrinsic Test of Unlearning Using Parametric Knowledge TracesYihuai Hong, Lei Yu, Haiqin Yang, Shauli Ravfogel 等EMNLP 2025 · 被引用 1 次
- Large Language Model Unlearning for Source CodeXue Jiang, Yihong Dong, Huangzhao Zhang, Tangxinyu Wang 等AAAI 2026
- LLM Unlearning Should Be Form-IndependentXiaotian Ye, Mengqi Zhang, Shu WuS&P 2026 · 被引用 3 次
- OBLIVIATE: Robust and Practical Machine Unlearning for Large Language ModelsXiaoyu Xu, Minxin Du, Qingqing Ye, Haibo HuEMNLP 2025 · 被引用 1 次
- Elastic Robust Unlearning of Specific Knowledge in Large Language ModelsYize Sui, Jing Ren, Wenjing Yang, Ruochun Jin 等NeurIPS 2025 · 被引用 1 次
