Direct Unlearning Optimization for Robust and Safe Text-to-Image Models
Yong-Hyun Park, Sangdoo Yun, Jin-Hwa Kim, Junho Kim, Geonhui Jang, Yonghyun Jeong, Junghyo Jo, Gayoung Lee
Abstract
Recent advancements in text-to-image (T2I) models have unlocked a wide range of applications but also present significant risks, particularly in their potential to generate unsafe content. To mitigate this issue, researchers have developed unlearning techniques to remove the model's ability to generate potentially harmful content. However, these methods are easily bypassed by adversarial attacks, making them unreliable for ensuring the safety of generated images. In this paper, we propose Direct Unlearning Optimization (DUO), a novel framework for removing Not Safe For Work (NSFW) content from T2I models while preserving their performance on unrelated topics. DUO employs a preference optimization approach using curated paired image data, ensuring that the model learns to remove unsafe visual concepts while retaining unrelated features. Furthermore, we introduce an output-preserving regularization term to maintain the model's generative capabilities on safe content. Extensive experiments demonstrate that DUO can robustly defend against various state-of-the-art red teaming methods without significant performance degradation on unrelated topics, as measured by FID and CLIP scores. Our work contributes to the development of safer and more reliable T2I models, paving the way for their responsible deployment in both closed-source and open-source scenarios.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d2eb895d-5e81-4732-aafb-3e1e7f48f6b8Cited by top-tier papers25
- CURE: Concept Unlearning via Orthogonal Representation Editing in Diffusion ModelsShristi Das Biswas, Arani Roy, Kaushik RoyNeurIPS 2025 · 30 citations
- EraseFlow: Learning Concept Erasure Policies via GFlowNet-Driven AlignmentNaga Sai Abhiram Kusumba, Maitreya Patel, Kyle Min, Changhoon Kim et al.NeurIPS 2025 · 10 citations
- Image Can Bring Your Memory Back: A Novel Multi-Modal Guided Attack against Image Generation Model UnlearningRenyang Liu, Guanlin Li, Tianwei Zhang, See-Kiong NgICLR 2026 · 9 citations
- Training-Free Safe Text Embedding Guidance for Text-to-Image Diffusion ModelsByeonghu Na, Mina Kang, Jiseok Kwak, Minsang Park et al.NeurIPS 2025 · 8 citations
- Red-Teaming Text-to-Image Systems by Rule-based Preference ModelingYichuan Cao, Yibo Miao, Xiao-Shan Gao, Yinpeng DongNeurIPS 2025 · 8 citations
Builds on30
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
Related papers
- ForceForget: Reinforcement Concept Removal for Enhancing Safety in Text-to-Image ModelsDong Han, Yong LiICML 2026
- SafeGen: Mitigating Sexually Explicit Content Generation in Text-to-Image ModelsXinfeng Li, Yuchen Yang, Jiangyi Deng, Chen Yan et al.CCS 2024 · 8 citations
- AdvI2I: Adversarial Image Attack on Image-to-Image Diffusion ModelsYaopei Zeng, Yuanpu Cao, Bochuan Cao, Yurui Chang et al.ICML 2025
- TarPro: Targeted Protection Against Malicious Image EditingKaixin Shen, Ruijie Quan, Jiaxu Miao, Jun XiaoAAAI 2026
- ART: Automatic Red-teaming for Text-to-Image Models to Protect Benign UsersGuanlin Li, Kangjie Chen, Shudong Zhang, Jie Zhang et al.NeurIPS 2024 · 39 citations
