CRISP: Persistent Concept Unlearning via Sparse Autoencoders
Tomer Ashuach, Dana Arad, Aaron Mueller, Martin Tutek, Yonatan Belinkov
Abstract
As large language models (LLMs) are increasingly deployed in real-world applications, the need to selectively remove unwanted knowledge while preserving model utility has become paramount. Recent work has explored sparse autoencoders (SAEs) to perform precise interventions on monosemantic features. However, most SAE-based methods operate at inference time, which does not create persistent changes in the model's parameters. Such interventions can be bypassed or reversed by malicious actors with parameter access. We introduce CRISP, a parameter-efficient method for persistent concept unlearning using SAEs. CRISP automatically identifies salient SAE features across multiple layers and suppresses their activations. We experiment with two LLMs and show that our method outperforms prior approaches on safety-critical unlearning tasks from the WMDP benchmark, successfully removing harmful knowledge while preserving general and in-domain capabilities. Featurelevel analysis reveals that CRISP achieves semantically coherent separation between target and benign concepts, allowing precise suppression of the target features. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a370e283-964e-4361-a2a9-dcd1bfd5a780Cited by top-tier papers2
- Can SAEs reveal and mitigate racial biases of LLMs in healthcare?Hiba Ahsan, Byron C. WallaceICLR 2026 · 1 citation
- Latent Agents: A Post-Training Procedure for Internalized Multi-Agent DebateJohn Seon Keun Yi, Aaron Mueller, Dokyun LeeACL 2026 · 1 citation
Builds on6
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Transformer Feed-Forward Layers Build Predictions by Promoting Concepts in the Vocabulary SpaceMor Geva, Avi Caciularu, Kevin Ro Wang, Yoav GoldbergEMNLP 2022 · 92 citations
- Erasing Conceptual Knowledge from Language ModelsRohit Gandikota, Sheridan Feucht, Samuel Marks, David BauNeurIPS 2025 · 35 citations
- Towards More Practical Threat Models in Artificial Intelligence SecurityKathrin Grosse, Lukas Bieringer, Tarek R. Besold, Alexandre AlahiUSENIX Security 2024 · 27 citations
- Precise In-Parameter Concept Erasure in Large Language ModelsYoav Gur-Arieh, Clara Suslik, Yihuai Hong, Fazl Barez et al.EMNLP 2025 · 10 citations
Related papers
- SAUCE: Selective Concept Unlearning in Vision-Language Models with Sparse AutoencodersJiahui Geng, Qing LiICCV 2025 · 12 citations
- Large Language Model Unlearning via Embedding-Corrupted PromptsChris Yuhao Liu, Yaxuan Wang, Jeffrey Flanigan, Yang LiuNeurIPS 2024 · 138 citations
- Model Unlearning via Sparse Autoencoder Subspace Guided ProjectionsXu Wang, Zihao Li, Benyou Wang, Yan Hu et al.EMNLP 2025 · 9 citations
- Adaptive Localization of Knowledge Negation for Continual LLM UnlearningAbudukelimu Wuerkaixi, Qizhou Wang, Sen Cui, Wutong Xu et al.ICML 2025
- Reinforcement Unlearning via Group Relative Policy OptimizationEfstratios Zaradoukas, Bardh Prenkaj, Gjergji KasneciICLR 2026 · 4 citations
