ConceptPrune: Concept Editing in Diffusion Models via Skilled Neuron Pruning
Ruchika Chavhan, Da Li, Timothy M. Hospedales
摘要
While large-scale text-to-image diffusion models have demonstrated impressive image-generation capabilities, there are significant concerns about their potential misuse for generating unsafe content, violating copyright, and perpetuating societal biases. Recently, the text-to-image generation community has begun addressing these concerns by editing or unlearning undesired concepts from pre-trained models. However, these methods often involve data-intensive and inefficient fine-tuning or utilize various forms of token remapping, rendering them susceptible to adversarial jailbreaks. In this paper, we present a simple and effective training-free approach, ConceptPrune, wherein we first identify critical regions within pretrained models responsible for generating undesirable concepts, thereby facilitating straightforward concept unlearning via weight pruning. Experiments across a range of concepts including artistic styles, nudity, object erasure, and gender debiasing demonstrate that target concepts can be efficiently erased by pruning a tiny fraction, approximately 0.12% of total weights, enabling multi-concept erasure and robustness against various white-box and black-box adversarial attacks. Our code is available at https://github.com/ruchikachavhan/concept-prune.git Introduction In recent years, text-to-image generation has witnessed significant advances driven by the development and adoption of diffusion models (DMs) [24, 43, 45, 46, 35, 60, 33, 39] across industries and realworld scenarios. However, this swift advancement presents a substantial risk. Diffusion models can threaten artists' livelihoods through style replication [11] , generate convincing deepfakes and NSFW content [40, 14] , and perpetuate societal biases [32] . The risks associated with large-scale textto-image models arise from billion-sized web-scraped datasets used in training, comprising public datasets like LAION [48], COYO [4], and CC12M [5] , that often lack human-level quality assurance. A simplistic and naive solution to mitigate these risks involves fine-tuning the model on datasets without this undesired content; however, this approach can prove to be highly compute-expensive. Several efforts addressing the risks of diffusion models have been made from the perspective of Concept Editing [26, 18, 19, 58, 36] and Model Unlearning (MU) [23, 65, 30, 56, 12] , both aimed at eliminating undesired prompts, albeit with differing objectives. Concept editing methods seek to eliminate undesired prompts by aligning latent representations of the target concept with a concept to be retained, via methods such as maximizing similarity [26, 18] and token remapping [58, 19] . Conversely, Model Unlearning formulates an objective that penalizes forgetting desired concepts while promoting the elimination of undesired ones, but this requires expensive computations and fine-tuning. Moreover, as most concept editing approaches rely on some form of token blacklisting or resteering [58] , adversarial attacks based on textual inversion [61, 38, 57, 53] have demonstrated the
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper21
- Unveiling Concept Attribution in Diffusion ModelsNguyen Hung-Quang, Hoang Phan, Khoa D. DoanNeurIPS 2025 · 被引用 13 次
- Image Can Bring Your Memory Back: A Novel Multi-Modal Guided Attack against Image Generation Model UnlearningRenyang Liu, Guanlin Li, Tianwei Zhang, See-Kiong NgICLR 2026 · 被引用 9 次
- Mass Concept Erasure in Diffusion Models with Concept HierarchyJiahang Tu, Ye Li, Yiming Wu, Hanbin Zhao 等AAAI 2026 · 被引用 7 次
- SAEmnesia: Erasing Concepts in Diffusion Models with Supervised Sparse AutoencodersEnrico Cassano, Riccardo Renzulli, Marco Nurisso, Mirko Zaffaroni 等ICML 2026 · 被引用 7 次
- Closing the Safety Gap: Surgical Concept Erasure in Visual Autoregressive ModelsXinhao Zhong, Yimin Zhou, Zhiqi Zhang, Junhao Li 等ICLR 2026 · 被引用 7 次
它引用的顶会 Paper26
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
相关 Paper
- Defensive Unlearning with Adversarial Training for Robust Concept Erasure in Diffusion ModelsYimeng Zhang, Xin Chen, Jinghan Jia, Yihua Zhang 等NeurIPS 2024 · 被引用 200 次
- Growth Inhibitors for Suppressing Inappropriate Image Concepts in Diffusion ModelsDie Chen, Zhiwen Li, Mingyuan Fan, Cen Chen 等ICLR 2025
- Erasing Concepts from Diffusion ModelsRohit Gandikota, Joanna Materzynska, Jaden Fiotto-Kaufman, David BauICCV 2023 · 被引用 536 次
- Ablating Concepts in Text-to-Image Diffusion ModelsNupur Kumari, Bingliang Zhang, Sheng-Yu Wang, Eli Shechtman 等ICCV 2023 · 被引用 327 次
- Roots Beneath the Cut: Uncovering the Risk of Concept Revival in Pruning-Based Unlearning for Diffusion ModelsCi Zhang, Zhaojun Ding, Chence Yang, Jun Liu 等CVPR 2026 · 被引用 1 次
