Precise In-Parameter Concept Erasure in Large Language Models
Yoav Gur-Arieh, Clara Suslik, Yihuai Hong, Fazl Barez, Mor Geva
Abstract
Large language models (LLMs) often acquire knowledge during pretraining that is undesirable in downstream deployments, e.g., sensitive information or copyrighted content. Existing approaches for removing such knowledge rely on fine-tuning, training low-rank adapters or fact-level editing, but these are either too coarse, too shallow, or ineffective. In this work, we propose PISCES (Precise In-parameter Suppression for Concept EraSure), a novel framework for precisely erasing entire concepts from model parameters by directly editing directions that encode them in parameter space. PISCES uses a disentangler model to decompose MLP vectors into interpretable features, identifies those associated with a target concept using automated interpretability techniques, and removes them from model parameters. Experiments on Gemma 2 and Llama 3.1 over various concepts show that PISCES achieves modest gains in efficacy over leading erasure methods, reducing accuracy on the target concept to as low as 7.7%, while dramatically improving erasure specificity (by up to 31%) and robustness (by up to 38%). Overall, these results demonstrate that feature-based in-parameter editing enables a more precise and reliable approach for removing conceptual knowledge in language models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 79d31265-ba38-4c4d-919e-996967dca2a4Cited by top-tier papers4
- Measuring Chain of Thought Faithfulness by Unlearning Reasoning StepsMartin Tutek, Fateme Hashemi Chaleshtori, Ana Marasovic, Yonatan BelinkovEMNLP 2025 · 37 citations
- Watch the Weights: Unsupervised monitoring and control of fine-tuned LLMsZiqian Zhong, Aditi RaghunathanICLR 2026 · 7 citations
- CRISP: Persistent Concept Unlearning via Sparse AutoencodersTomer Ashuach, Dana Arad, Aaron Mueller, Martin Tutek et al.ACL 2026 · 7 citations
- Safety Anchor: Defending Harmful Fine-tuning via Geometric BottlenecksGuoxin Lu, Letian Sha, Qing Wang, Peijie Sun et al.ICML 2026
Builds on26
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 3,415 citations
- Refusal in Language Models Is Mediated by a Single DirectionAndy Arditi, Oscar Obeso, Aaquib Syed, Daniel Paleka et al.NeurIPS 2024 · 1,166 citations
- Sparse Autoencoders Find Highly Interpretable Features in Language ModelsRobert Huben, Hoagy Cunningham, Logan Riggs Smith, Aidan Ewart et al.ICLR 2024 · 1,072 citations
Related papers
- LEACE: Perfect linear concept erasure in closed formNora Belrose, David Schneider-Joseph, Shauli Ravfogel, Ryan Cotterell et al.NeurIPS 2023 · 305 citations
- Prototype-Guided Concept Erasure in Diffusion ModelsYuze Cai, Jiahao Lu, Hongxiang Shi, Yichao Zhou et al.CVPR 2026 · 3 citations
- Disentangling Knowledge Representations for Large Language Model EditingMengqi Zhang, Zisheng Zhou, Xiaotian Ye, Qiang Liu et al.ICLR 2026 · 6 citations
- EMMA: Concept Erasure Benchmark with Comprehensive Semantic Metrics and Diverse CategoriesLu Wei, Yuta Nakashima, Noa GarciaCVPR 2026 · 4 citations
- LLM-Eraser: Optimizing Large Language Model Unlearning through Selective PruningShengming Zhang, Le Zhang, Jingbo Zhou, Zhi Zheng et al.KDD 2025 · 2 citations
