Boosting Alignment for Post-Unlearning Text-to-Image Generative Models
Myeongseob Ko, Henry Li, Zhun Wang, Jonathan Patsenker, Jiachen T. Wang, Qinbin Li, Ming Jin, Dawn Song, Ruoxi Jia
Abstract
Large-scale generative models have shown impressive image-generation capabilities, propelled by massive data. However, this often inadvertently leads to the generation of harmful or inappropriate content and raises copyright concerns. Driven by these concerns, machine unlearning has become crucial to effectively purge undesirable knowledge from models. While existing literature has studied various unlearning techniques, these often suffer from either poor unlearning quality or degradation in text-image alignment after unlearning, due to the competitive nature of these objectives. To address these challenges, we propose a framework that seeks an optimal model update at each unlearning iteration, ensuring monotonic improvement on both objectives. We further derive the characterization of such an update. In addition, we design procedures to strategically diversify the unlearning and remaining datasets to boost performance improvement. Our evaluation demonstrates that our method effectively removes target classes from recent diffusion-based generative models and concepts from stable diffusion models while maintaining close alignment with the models' original trained states, thus outperforming state-of-the-art baselines. Our code will be made available at https://github.com/reds-lab/Restricted_gradient_diversity_unlearning.git.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 549d7f5d-3ade-401c-bbef-307f146714e1Cited by top-tier papers8
- Unlearning Concepts in Diffusion Model via Concept Domain Correction and Concept Preserving GradientYongliang Wu, Shiji Zhou, Mingzhuo Yang, Lianzhe Wang et al.AAAI 2025 · 69 citations
- Learning to Unlearn While Retaining: Combating Gradient Conflicts in Machine UnlearningGaurav Patel, Qiang QiuICCV 2025 · 21 citations
- Sharpness-Aware Machine UnlearningHaoran Tang, Rajiv KhannaICLR 2026 · 10 citations
- Unlearning-Aware MinimizationHoki Kim, Keonwoo Kim, Sungwon Chae, Sangwon YoonNeurIPS 2025 · 7 citations
- Probing Hidden Knowledge Holes in Unlearned LLMsMyeongseob Ko, Hoang Anh Just, Charles Fleming, Ming Jin et al.NeurIPS 2025 · 3 citations
Builds on18
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Elucidating the Design Space of Diffusion-Based Generative ModelsTero Karras, Miika Aittala, Timo Aila, Samuli LaineNeurIPS 2022 · 3,959 citations
- Gradient Surgery for Multi-Task LearningTianhe Yu, Saurabh Kumar, Abhishek Gupta, Sergey Levine et al.NeurIPS 2020 · 2,261 citations
Related papers
- Meta-Unlearning on Diffusion Models: Preventing Relearning Unlearned ConceptsHongcheng Gao, Tianyu Pang, Chao Du, Taihang Hu et al.ICCV 2025 · 4 citations
- Sculpting Memory: Multi-Concept Forgetting in Diffusion Models via Dynamic Mask and Concept-Aware OptimizationGen Li, Yang Xiao, Jie Ji, Kaiyuan Deng et al.ICCV 2025 · 1 citation
- Ablating Concepts in Text-to-Image Diffusion ModelsNupur Kumari, Bingliang Zhang, Sheng-Yu Wang, Eli Shechtman et al.ICCV 2023 · 327 citations
- The Illusion of Unlearning: The Unstable Nature of Machine Unlearning in Text-to-Image Diffusion ModelsNaveen George, Karthik Nandan Dasaraju, Rutheesh Reddy Chittepu, Konda Reddy MopuriCVPR 2025
- Preference-Calibrated Optimization with Score-Level Distribution Alignment for Text-to-Image Diffusion Model UnlearningXiuyuan Wang, Weiming Liu, Hongyu Cai, Xin Gao et al.ICML 2026
