ConceptPrune: Concept Editing in Diffusion Models via Skilled Neuron Pruning
Ruchika Chavhan, Da Li, Timothy M. Hospedales
Abstract
While large-scale text-to-image diffusion models have demonstrated impressive image-generation capabilities, there are significant concerns about their potential misuse for generating unsafe content, violating copyright, and perpetuating societal biases. Recently, the text-to-image generation community has begun addressing these concerns by editing or unlearning undesired concepts from pre-trained models. However, these methods often involve data-intensive and inefficient fine-tuning or utilize various forms of token remapping, rendering them susceptible to adversarial jailbreaks. In this paper, we present a simple and effective training-free approach, ConceptPrune, wherein we first identify critical regions within pretrained models responsible for generating undesirable concepts, thereby facilitating straightforward concept unlearning via weight pruning. Experiments across a range of concepts including artistic styles, nudity, object erasure, and gender debiasing demonstrate that target concepts can be efficiently erased by pruning a tiny fraction, approximately 0.12% of total weights, enabling multi-concept erasure and robustness against various white-box and black-box adversarial attacks. Our code is available at https://github.com/ruchikachavhan/concept-prune.git Introduction In recent years, text-to-image generation has witnessed significant advances driven by the development and adoption of diffusion models (DMs) [24, 43, 45, 46, 35, 60, 33, 39] across industries and realworld scenarios. However, this swift advancement presents a substantial risk. Diffusion models can threaten artists' livelihoods through style replication [11] , generate convincing deepfakes and NSFW content [40, 14] , and perpetuate societal biases [32] . The risks associated with large-scale textto-image models arise from billion-sized web-scraped datasets used in training, comprising public datasets like LAION [48], COYO [4], and CC12M [5] , that often lack human-level quality assurance. A simplistic and naive solution to mitigate these risks involves fine-tuning the model on datasets without this undesired content; however, this approach can prove to be highly compute-expensive. Several efforts addressing the risks of diffusion models have been made from the perspective of Concept Editing [26, 18, 19, 58, 36] and Model Unlearning (MU) [23, 65, 30, 56, 12] , both aimed at eliminating undesired prompts, albeit with differing objectives. Concept editing methods seek to eliminate undesired prompts by aligning latent representations of the target concept with a concept to be retained, via methods such as maximizing similarity [26, 18] and token remapping [58, 19] . Conversely, Model Unlearning formulates an objective that penalizes forgetting desired concepts while promoting the elimination of undesired ones, but this requires expensive computations and fine-tuning. Moreover, as most concept editing approaches rely on some form of token blacklisting or resteering [58] , adversarial attacks based on textual inversion [61, 38, 57, 53] have demonstrated the
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7de874fe-80be-477a-bede-c967376cbf7aCited by top-tier papers21
- Unveiling Concept Attribution in Diffusion ModelsNguyen Hung-Quang, Hoang Phan, Khoa D. DoanNeurIPS 2025 · 13 citations
- Image Can Bring Your Memory Back: A Novel Multi-Modal Guided Attack against Image Generation Model UnlearningRenyang Liu, Guanlin Li, Tianwei Zhang, See-Kiong NgICLR 2026 · 9 citations
- Mass Concept Erasure in Diffusion Models with Concept HierarchyJiahang Tu, Ye Li, Yiming Wu, Hanbin Zhao et al.AAAI 2026 · 7 citations
- SAEmnesia: Erasing Concepts in Diffusion Models with Supervised Sparse AutoencodersEnrico Cassano, Riccardo Renzulli, Marco Nurisso, Mirko Zaffaroni et al.ICML 2026 · 7 citations
- Closing the Safety Gap: Surgical Concept Erasure in Visual Autoregressive ModelsXinhao Zhong, Yimin Zhou, Zhiqi Zhang, Junhao Li et al.ICLR 2026 · 7 citations
Builds on26
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
Related papers
- Defensive Unlearning with Adversarial Training for Robust Concept Erasure in Diffusion ModelsYimeng Zhang, Xin Chen, Jinghan Jia, Yihua Zhang et al.NeurIPS 2024 · 200 citations
- Growth Inhibitors for Suppressing Inappropriate Image Concepts in Diffusion ModelsDie Chen, Zhiwen Li, Mingyuan Fan, Cen Chen et al.ICLR 2025
- Erasing Concepts from Diffusion ModelsRohit Gandikota, Joanna Materzynska, Jaden Fiotto-Kaufman, David BauICCV 2023 · 536 citations
- Ablating Concepts in Text-to-Image Diffusion ModelsNupur Kumari, Bingliang Zhang, Sheng-Yu Wang, Eli Shechtman et al.ICCV 2023 · 327 citations
- Roots Beneath the Cut: Uncovering the Risk of Concept Revival in Pruning-Based Unlearning for Diffusion ModelsCi Zhang, Zhaojun Ding, Chence Yang, Jun Liu et al.CVPR 2026 · 1 citation
