Fantastic Targets for Concept Erasure in Diffusion Models and Where To Find Them
Anh Tuan Bui, Thuy-Trang Vu, Long Tung Vuong, Trung Le, Paul Montague, Tamas Abraham, Junae Kim, Dinh Phung
Abstract
Concept erasure has emerged as a promising technique for mitigating the risk of harmful content generation in diffusion models by selectively unlearning undesirable concepts. The common principle of previous works to remove a specific concept is to map it to a fixed generic concept, such as a neutral concept or just an empty text prompt. In this paper, we demonstrate that this fixed-target strategy is suboptimal, as it fails to account for the impact of erasing one concept on the others. To address this limitation, we model the concept space as a graph and empirically analyze the effects of erasing one concept on the remaining concepts. Our analysis uncovers intriguing geometric properties of the concept space, where the influence of erasing a concept is confined to a local region. Building on this insight, we propose the Adaptive Guided Erasure (AGE) method, which dynamically selects optimal target concepts tailored to each undesirable concept, minimizing unintended side effects. Experimental results show that AGE significantly outperforms state-of-the-art erasure methods on preserving unrelated concepts while maintaining effective erasure performance. Our code is published at https://github.com/tuananhbui89/Adaptive-Guided-Erasure.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4fa8efef-e9a8-47e5-8d00-3c73e92e6dc6Cited by top-tier papers17
- EraseFlow: Learning Concept Erasure Policies via GFlowNet-Driven AlignmentNaga Sai Abhiram Kusumba, Maitreya Patel, Kyle Min, Changhoon Kim et al.NeurIPS 2025 · 10 citations
- Continual Unlearning for Text-to-Image Diffusion Models: A Regularization PerspectiveJustin Lee, Zheda Mai, Jinsu Yoo, Chongyu Fan et al.ICLR 2026 · 9 citations
- Closing the Safety Gap: Surgical Concept Erasure in Visual Autoregressive ModelsXinhao Zhong, Yimin Zhou, Zhiqi Zhang, Junhao Li et al.ICLR 2026 · 7 citations
- Co-occurring Associated REtained concepts in Diffusion UnlearningMiso Kim, Georu Lee, Yunji Kim, Hoki Kim et al.ICLR 2026 · 6 citations
- AEGIS: Adversarial Target-Guided Retention-Data-Free Robust Concept Erasure from Diffusion ModelsFengpeng Li, Kemou Li, Qizhou Wang, Bo Han et al.ICLR 2026 · 6 citations
Builds on13
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
- Erasing Concepts from Diffusion ModelsRohit Gandikota, Joanna Materzynska, Jaden Fiotto-Kaufman, David BauICCV 2023 · 536 citations
- An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual InversionRinon Gal, Yuval Alaluf, Yuval Atzmon, Or Patashnik et al.ICLR 2023 · 464 citations
Related papers
- Erasing Undesirable Concepts in Diffusion Models with Adversarial PreservationAnh Bui, Tung-Long Vuong, Khanh Doan, Trung Le et al.NeurIPS 2024 · 55 citations
- GrOCE : Graph-Guided Online Concept Erasure for Text-to-Image Diffusion ModelsNing Han, Zhenyu Ge, Feng Han, Yuhua Sun et al.CVPR 2026 · 3 citations
- One-dimensional Adapter to Rule Them All: Concepts, Diffusion Models and Erasing ApplicationsMengyao Lyu, Yuhong Yang, Haiwen Hong, Hui Chen et al.CVPR 2024 · 16 citations
- TRCE: Towards Reliable Malicious Concept Erasure in Text-to-Image Diffusion ModelsRuidong Chen, Honglin Guo, Lanjun Wang, Chenyu Zhang et al.ICCV 2025 · 17 citations
- When Are Concepts Erased From Diffusion Models?Kevin Lu, Nicky Kriplani, Rohit Gandikota, Minh Pham et al.NeurIPS 2025 · 21 citations
