CRISP: Persistent Concept Unlearning via Sparse Autoencoders
Tomer Ashuach, Dana Arad, Aaron Mueller, Martin Tutek, Yonatan Belinkov
摘要
As large language models (LLMs) are increasingly deployed in real-world applications, the need to selectively remove unwanted knowledge while preserving model utility has become paramount. Recent work has explored sparse autoencoders (SAEs) to perform precise interventions on monosemantic features. However, most SAE-based methods operate at inference time, which does not create persistent changes in the model's parameters. Such interventions can be bypassed or reversed by malicious actors with parameter access. We introduce CRISP, a parameter-efficient method for persistent concept unlearning using SAEs. CRISP automatically identifies salient SAE features across multiple layers and suppresses their activations. We experiment with two LLMs and show that our method outperforms prior approaches on safety-critical unlearning tasks from the WMDP benchmark, successfully removing harmful knowledge while preserving general and in-domain capabilities. Featurelevel analysis reveals that CRISP achieves semantically coherent separation between target and benign concepts, allowing precise suppression of the target features. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Can SAEs reveal and mitigate racial biases of LLMs in healthcare?Hiba Ahsan, Byron C. WallaceICLR 2026 · 被引用 1 次
- Latent Agents: A Post-Training Procedure for Internalized Multi-Agent DebateJohn Seon Keun Yi, Aaron Mueller, Dokyun LeeACL 2026 · 被引用 1 次
它引用的顶会 Paper6
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- Transformer Feed-Forward Layers Build Predictions by Promoting Concepts in the Vocabulary SpaceMor Geva, Avi Caciularu, Kevin Ro Wang, Yoav GoldbergEMNLP 2022 · 被引用 92 次
- Erasing Conceptual Knowledge from Language ModelsRohit Gandikota, Sheridan Feucht, Samuel Marks, David BauNeurIPS 2025 · 被引用 35 次
- Towards More Practical Threat Models in Artificial Intelligence SecurityKathrin Grosse, Lukas Bieringer, Tarek R. Besold, Alexandre AlahiUSENIX Security 2024 · 被引用 27 次
- Precise In-Parameter Concept Erasure in Large Language ModelsYoav Gur-Arieh, Clara Suslik, Yihuai Hong, Fazl Barez 等EMNLP 2025 · 被引用 10 次
相关 Paper
- SAUCE: Selective Concept Unlearning in Vision-Language Models with Sparse AutoencodersJiahui Geng, Qing LiICCV 2025 · 被引用 12 次
- Large Language Model Unlearning via Embedding-Corrupted PromptsChris Yuhao Liu, Yaxuan Wang, Jeffrey Flanigan, Yang LiuNeurIPS 2024 · 被引用 138 次
- Model Unlearning via Sparse Autoencoder Subspace Guided ProjectionsXu Wang, Zihao Li, Benyou Wang, Yan Hu 等EMNLP 2025 · 被引用 9 次
- Adaptive Localization of Knowledge Negation for Continual LLM UnlearningAbudukelimu Wuerkaixi, Qizhou Wang, Sen Cui, Wutong Xu 等ICML 2025
- Reinforcement Unlearning via Group Relative Policy OptimizationEfstratios Zaradoukas, Bardh Prenkaj, Gjergji KasneciICLR 2026 · 被引用 4 次
