Fantastic Targets for Concept Erasure in Diffusion Models and Where To Find Them
Anh Tuan Bui, Thuy-Trang Vu, Long Tung Vuong, Trung Le, Paul Montague, Tamas Abraham, Junae Kim, Dinh Phung
摘要
Concept erasure has emerged as a promising technique for mitigating the risk of harmful content generation in diffusion models by selectively unlearning undesirable concepts. The common principle of previous works to remove a specific concept is to map it to a fixed generic concept, such as a neutral concept or just an empty text prompt. In this paper, we demonstrate that this fixed-target strategy is suboptimal, as it fails to account for the impact of erasing one concept on the others. To address this limitation, we model the concept space as a graph and empirically analyze the effects of erasing one concept on the remaining concepts. Our analysis uncovers intriguing geometric properties of the concept space, where the influence of erasing a concept is confined to a local region. Building on this insight, we propose the Adaptive Guided Erasure (AGE) method, which dynamically selects optimal target concepts tailored to each undesirable concept, minimizing unintended side effects. Experimental results show that AGE significantly outperforms state-of-the-art erasure methods on preserving unrelated concepts while maintaining effective erasure performance. Our code is published at https://github.com/tuananhbui89/Adaptive-Guided-Erasure.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- EraseFlow: Learning Concept Erasure Policies via GFlowNet-Driven AlignmentNaga Sai Abhiram Kusumba, Maitreya Patel, Kyle Min, Changhoon Kim 等NeurIPS 2025 · 被引用 10 次
- Continual Unlearning for Text-to-Image Diffusion Models: A Regularization PerspectiveJustin Lee, Zheda Mai, Jinsu Yoo, Chongyu Fan 等ICLR 2026 · 被引用 9 次
- Closing the Safety Gap: Surgical Concept Erasure in Visual Autoregressive ModelsXinhao Zhong, Yimin Zhou, Zhiqi Zhang, Junhao Li 等ICLR 2026 · 被引用 7 次
- Co-occurring Associated REtained concepts in Diffusion UnlearningMiso Kim, Georu Lee, Yunji Kim, Hoki Kim 等ICLR 2026 · 被引用 6 次
- AEGIS: Adversarial Target-Guided Retention-Data-Free Robust Concept Erasure from Diffusion ModelsFengpeng Li, Kemou Li, Qizhou Wang, Bo Han 等ICLR 2026 · 被引用 6 次
它引用的顶会 Paper13
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray 等ICML 2021 · 被引用 6,356 次
- Erasing Concepts from Diffusion ModelsRohit Gandikota, Joanna Materzynska, Jaden Fiotto-Kaufman, David BauICCV 2023 · 被引用 536 次
- An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual InversionRinon Gal, Yuval Alaluf, Yuval Atzmon, Or Patashnik 等ICLR 2023 · 被引用 464 次
相关 Paper
- Erasing Undesirable Concepts in Diffusion Models with Adversarial PreservationAnh Bui, Tung-Long Vuong, Khanh Doan, Trung Le 等NeurIPS 2024 · 被引用 55 次
- GrOCE : Graph-Guided Online Concept Erasure for Text-to-Image Diffusion ModelsNing Han, Zhenyu Ge, Feng Han, Yuhua Sun 等CVPR 2026 · 被引用 3 次
- One-dimensional Adapter to Rule Them All: Concepts, Diffusion Models and Erasing ApplicationsMengyao Lyu, Yuhong Yang, Haiwen Hong, Hui Chen 等CVPR 2024 · 被引用 16 次
- TRCE: Towards Reliable Malicious Concept Erasure in Text-to-Image Diffusion ModelsRuidong Chen, Honglin Guo, Lanjun Wang, Chenyu Zhang 等ICCV 2025 · 被引用 17 次
- When Are Concepts Erased From Diffusion Models?Kevin Lu, Nicky Kriplani, Rohit Gandikota, Minh Pham 等NeurIPS 2025 · 被引用 21 次
