Meta-Unlearning on Diffusion Models: Preventing Relearning Unlearned Concepts
Hongcheng Gao, Tianyu Pang, Chao Du, Taihang Hu, Zhijie Deng, Min Lin
Abstract
With the rapid progress of diffusion-based content generation, significant efforts are being made to unlearn harmful or copyrighted concepts from pretrained diffusion models (DMs) to prevent potential model misuse. However, it is observed that even when DMs are properly unlearned before release, malicious finetuning can compromise this process, causing DMs to relearn the unlearned concepts. This occurs partly because certain benign concepts (e.g., "skin") retained in DMs are related to the unlearned ones (e.g., "nudity"), facilitating their relearning via finetuning. To address this, we propose meta-unlearning on DMs. Intuitively, a meta-unlearned DM should behave like an unlearned DM when used as is; moreover, if the meta-unlearned DM undergoes malicious finetuning on unlearned concepts, the related benign concepts retained within it will be triggered to self-destruct, hindering the relearning of unlearned concepts. Our meta-unlearning framework is compatible with most existing unlearning methods, requiring only the addition of an easy-to-implement meta objective. We validate our approach through empirical experiments on meta-unlearning concepts from Stable Diffusion models (SD-v1-4 and SDXL), supported by extensive ablation studies. Our code is available at https://github.com/sail-sg/Meta-Unlearning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6d61bfd1-a289-4fe5-9278-1eefb6a3b672Cited by top-tier papers17
- Unlearning Concepts in Diffusion Model via Concept Domain Correction and Concept Preserving GradientYongliang Wu, Shiji Zhou, Mingzhuo Yang, Lianzhe Wang et al.AAAI 2025 · 69 citations
- CTRAP: Embedding Collapse Trap to Safeguard Large Language Models from Harmful Fine-TuningBiao Yi, Tiansheng Huang, Baolei Zhang, Tong Li et al.ACL 2026 · 15 citations
- Leveraging Catastrophic Forgetting to Develop Safe Diffusion Models against Malicious FinetuningJiadong Pan, Hongcheng Gao, Zongyu Wu, Taihang Hu et al.NeurIPS 2024 · 15 citations
- Towards Resilient Safety-driven Unlearning for Diffusion Models against Downstream Fine-tuningBoheng Li, Renjie Gu, Junjie Wang, Leyi Qi et al.NeurIPS 2025 · 15 citations
- Reminiscence Attack on Residuals: Exploiting Approximate Machine Unlearning for PrivacyYaxin Xiao, Qingqing Ye, Li Hu, Huadi Zheng et al.ICCV 2025 · 6 citations
Builds on29
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann et al.ICLR 2024 · 4,569 citations
- Score-Based Generative Modeling through Stochastic Differential EquationsYang Song, Jascha Sohl-Dickstein, Diederik P. Kingma, Abhishek Kumar et al.ICLR 2021 · 1,270 citations
- Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!Xiangyu Qi, Yi Zeng, Tinghao Xie, Pin-Yu Chen et al.ICLR 2024 · 1,104 citations
Related papers
- Boosting Alignment for Post-Unlearning Text-to-Image Generative ModelsMyeongseob Ko, Henry Li, Zhun Wang, Jonathan Patsenker et al.NeurIPS 2024 · 22 citations
- The Illusion of Unlearning: The Unstable Nature of Machine Unlearning in Text-to-Image Diffusion ModelsNaveen George, Karthik Nandan Dasaraju, Rutheesh Reddy Chittepu, Konda Reddy MopuriCVPR 2025
- Ablating Concepts in Text-to-Image Diffusion ModelsNupur Kumari, Bingliang Zhang, Sheng-Yu Wang, Eli Shechtman et al.ICCV 2023 · 327 citations
- Defensive Unlearning with Adversarial Training for Robust Concept Erasure in Diffusion ModelsYimeng Zhang, Xin Chen, Jinghan Jia, Yihua Zhang et al.NeurIPS 2024 · 200 citations
- Co-occurring Associated REtained concepts in Diffusion UnlearningMiso Kim, Georu Lee, Yunji Kim, Hoki Kim et al.ICLR 2026 · 6 citations
