Defensive Unlearning with Adversarial Training for Robust Concept Erasure in Diffusion Models
Yimeng Zhang, Xin Chen, Jinghan Jia, Yihua Zhang, Chongyu Fan, Jiancheng Liu, Mingyi Hong, Ke Ding, Sijia Liu
摘要
Diffusion models (DMs) have achieved remarkable success in text-to-image generation, but they also pose safety risks, such as the potential generation of harmful content and copyright violations. The techniques of machine unlearning, also known as concept erasing, have been developed to address these risks. However, these techniques remain vulnerable to adversarial prompt attacks, which can prompt DMs post-unlearning to regenerate undesired images containing concepts (such as nudity) meant to be erased. This work aims to enhance the robustness of concept erasing by integrating the principle of adversarial training (AT) into machine unlearning, resulting in the robust unlearning framework referred to as AdvUnlearn. However, achieving this effectively and efficiently is highly nontrivial. First, we find that a straightforward implementation of AT compromises DMs' image generation quality post-unlearning. To address this, we develop a utility-retaining regularization on an additional retain set, optimizing the trade-off between concept erasure robustness and model utility in AdvUnlearn. Moreover, we identify the text encoder as a more suitable module for robustification compared to UNet, ensuring unlearning effectiveness. And the acquired text encoder can serve as a plug-and-play robust unlearner for various DM types. Empirically, we perform extensive experiments to demonstrate the robustness advantage of AdvUnlearn across various DM unlearning scenarios, including the erasure of nudity, objects, and style concepts. In addition to robustness, AdvUnlearn also achieves a balanced tradeoff with model utility. To our knowledge, this is the first work to systematically explore robust DM unlearning through AT, setting it apart from existing methods that overlook robustness in concept erasing. Codes are available at: https://github.com/OPTML-Group/AdvUnlearn
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper68
- Unlearning Concepts in Diffusion Model via Concept Domain Correction and Concept Preserving GradientYongliang Wu, Shiji Zhou, Mingzhuo Yang, Lianzhe Wang 等AAAI 2025 · 被引用 69 次
- Erasing Undesirable Concepts in Diffusion Models with Adversarial PreservationAnh Bui, Tung-Long Vuong, Khanh Doan, Trung Le 等NeurIPS 2024 · 被引用 55 次
- SPEED: Scalable, Precise, and Efficient Concept Erasure for Diffusion ModelsOuxiang Li, Yuan Wang, Xinting Hu, Houcheng Jiang 等ICLR 2026 · 被引用 37 次
- Learning to Unlearn While Retaining: Combating Gradient Conflicts in Machine UnlearningGaurav Patel, Qiang QiuICCV 2025 · 被引用 21 次
- TRCE: Towards Reliable Malicious Concept Erasure in Text-to-Image Diffusion ModelsRuidong Chen, Honglin Guo, Lanjun Wang, Chenyu Zhang 等ICCV 2025 · 被引用 17 次
它引用的顶会 Paper46
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 被引用 9,786 次
- Improved Denoising Diffusion Probabilistic ModelsAlexander Quinn Nichol, Prafulla DhariwalICML 2021 · 被引用 5,234 次
相关 Paper
- Image Can Bring Your Memory Back: A Novel Multi-Modal Guided Attack against Image Generation Model UnlearningRenyang Liu, Guanlin Li, Tianwei Zhang, See-Kiong NgICLR 2026 · 被引用 9 次
- Adaptive Median Smoothing: Adversarial Defense for Unlearned Text-to-Image Diffusion Models at Inference TimeXiaoxuan Han, Songlin Yang, Wei Wang, Yang Li 等ICML 2025
- ConceptPrune: Concept Editing in Diffusion Models via Skilled Neuron PruningRuchika Chavhan, Da Li, Timothy M. HospedalesICLR 2025 · 被引用 2 次
- Meta-Unlearning on Diffusion Models: Preventing Relearning Unlearned ConceptsHongcheng Gao, Tianyu Pang, Chao Du, Taihang Hu 等ICCV 2025 · 被引用 4 次
- The Illusion of Unlearning: The Unstable Nature of Machine Unlearning in Text-to-Image Diffusion ModelsNaveen George, Karthik Nandan Dasaraju, Rutheesh Reddy Chittepu, Konda Reddy MopuriCVPR 2025
