Erasing Undesirable Concepts in Diffusion Models with Adversarial Preservation
Anh Bui, Tung-Long Vuong, Khanh Doan, Trung Le, Paul Montague, Tamas Abraham, Dinh Q. Phung
Abstract
Diffusion models excel at generating visually striking content from text but can inadvertently produce undesirable or harmful content when trained on unfiltered internet data. A practical solution is to selectively removing target concepts from the model, but this may impact the remaining concepts. Prior approaches have tried to balance this by introducing a loss term to preserve neutral content or a regularization term to minimize changes in the model parameters, yet resolving this trade-off remains challenging. In this work, we propose to identify and preserving concepts most affected by parameter changes, termed as adversarial concepts. This approach ensures stable erasure with minimal impact on the other concepts. We demonstrate the effectiveness of our method using the Stable Diffusion model, showing that it outperforms state-of-the-art erasure methods in eliminating unwanted content while maintaining the integrity of other unrelated elements. Our code is available at https://github.com/tuananhbui89/Erasing-Adversarial-Preservation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext aa84bf0e-680d-4caf-8a21-b7dafa97eff0Cited by top-tier papers23
- On Effects of Steering Latent Representation for Large Language Model UnlearningHuu-Tien Dang, Tin Pham, Hoang Thanh-Tung, Naoya InoueAAAI 2025 · 33 citations
- Continual Unlearning for Text-to-Image Diffusion Models: A Regularization PerspectiveJustin Lee, Zheda Mai, Jinsu Yoo, Chongyu Fan et al.ICLR 2026 · 9 citations
- Semantic Surgery: Zero-Shot Concept Erasure in Diffusion ModelsLexiang Xiong, Chengyu Liu, Jingwen Ye, Yan Liu et al.NeurIPS 2025 · 8 citations
- Closing the Safety Gap: Surgical Concept Erasure in Visual Autoregressive ModelsXinhao Zhong, Yimin Zhou, Zhiqi Zhang, Junhao Li et al.ICLR 2026 · 7 citations
- AEGIS: Adversarial Target-Guided Retention-Data-Free Robust Concept Erasure from Diffusion ModelsFengpeng Li, Kemou Li, Qizhou Wang, Bo Han et al.ICLR 2026 · 6 citations
Builds on18
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
- Erasing Concepts from Diffusion ModelsRohit Gandikota, Joanna Materzynska, Jaden Fiotto-Kaufman, David BauICCV 2023 · 536 citations
- An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual InversionRinon Gal, Yuval Alaluf, Yuval Atzmon, Or Patashnik et al.ICLR 2023 · 464 citations
Related papers
- Fantastic Targets for Concept Erasure in Diffusion Models and Where To Find ThemAnh Tuan Bui, Thuy-Trang Vu, Long Tung Vuong, Trung Le et al.ICLR 2025
- Ablating Concepts in Text-to-Image Diffusion ModelsNupur Kumari, Bingliang Zhang, Sheng-Yu Wang, Eli Shechtman et al.ICCV 2023 · 327 citations
- SDErasure: Concept-Specific Trajectory Shifting for Concept Erasure via Adaptive Diffusion ClassifierFengyuan Miao, Shancheng Fang, Lingyun Yu, Yadong Qu et al.ICLR 2026
- Concept Pinpoint Eraser for Text-to-image Diffusion Models via Residual Attention GateByung Hyun Lee, Sungjin Lim, Seunggyu Lee, Dong Un Kang et al.ICLR 2025
- Memories of Forgotten ConceptsMatan Rusanovsky, Shimon Malnick, Amir Jevnisek, Ohad Fried et al.CVPR 2025
