Circumventing Concept Erasure Methods For Text-To-Image Generative Models
Minh Pham, Kelly O. Marshall, Niv Cohen, Govind Mittal, Chinmay Hegde
摘要
Text-to-image generative models can produce photo-realistic images for an extremely broad range of concepts, and their usage has proliferated widely among the general public. On the flip side, these models have numerous drawbacks, including their potential to generate images featuring sexually explicit content, mirror artistic styles without permission, or even hallucinate (or deepfake) the likenesses of celebrities. Consequently, various methods have been proposed in order to"erase"sensitive concepts from text-to-image models. In this work, we examine five recently proposed concept erasure methods, and show that targeted concepts are not fully excised from any of these methods. Specifically, we leverage the existence of special learned word embeddings that can retrieve"erased"concepts from the sanitized models with no alterations to their weights. Our results highlight the brittleness of post hoc concept erasure methods, and call into question their use in the algorithmic toolkit for AI safety.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper37
- Defensive Unlearning with Adversarial Training for Robust Concept Erasure in Diffusion ModelsYimeng Zhang, Xin Chen, Jinghan Jia, Yihua Zhang 等NeurIPS 2024 · 被引用 200 次
- Direct Unlearning Optimization for Robust and Safe Text-to-Image ModelsYong-Hyun Park, Sangdoo Yun, Jin-Hwa Kim, Junho Kim 等NeurIPS 2024 · 被引用 60 次
- CURE: Concept Unlearning via Orthogonal Representation Editing in Diffusion ModelsShristi Das Biswas, Arani Roy, Kaushik RoyNeurIPS 2025 · 被引用 30 次
- Yuan: Yielding Unblemished Aesthetics Through a Unified Network for Visual Imperfections Removal in Generated ImagesZhenyu Yu, Chee Seng ChanAAAI 2025 · 被引用 22 次
- Training-Free Safe Denoisers for Safe Use of Diffusion ModelsMingyu Kim, Dongjun Kim, Amman Yusuf, Stefano Ermon 等NeurIPS 2025 · 被引用 21 次
它引用的顶会 Paper23
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
相关 Paper
- Prototype-Guided Concept Erasure in Diffusion ModelsYuze Cai, Jiahao Lu, Hongxiang Shi, Yichao Zhou 等CVPR 2026 · 被引用 3 次
- ForceForget: Reinforcement Concept Removal for Enhancing Safety in Text-to-Image ModelsDong Han, Yong LiICML 2026
- Erasing Concepts from Diffusion ModelsRohit Gandikota, Joanna Materzynska, Jaden Fiotto-Kaufman, David BauICCV 2023 · 被引用 536 次
- Beyond Text Prompts: Precise Concept Erasure through Text-Image CollaborationJun Li, Lizhi Xiong, Ziqiang Li, Weiwei Jiang 等CVPR 2026 · 被引用 1 次
- Erased but Not Forgotten: How Backdoors Compromise Concept ErasureTobias Braun, Jonas Henry Grebe, Patrick Mohr Gordillo, Marcus Rohrbach 等ICML 2026
