AlignGuard: Scalable Safety Alignment for Text-to-Image Generation
Runtao Liu, Chen I Chieh, Jindong Gu, Jipeng Zhang, Renjie Pi, Qifeng Chen, Philip Torr, Ashkan Khakzar, Fabio Pizzati
Abstract
Text-to-image (T2I) models are widespread, but their limited safety guardrails expose end users to harmful content and potentially allow for model misuse. Current safety measures are typically limited to text-based filtering or concept removal strategies, able to remove just a few concepts from the model's generative capabilities. In this work, we introduce AlignGuard, a method for safety alignment of T2I models. We enable the application of Direct Preference Optimization (DPO) for safety purposes in T2I models by synthetically generating a dataset of harmful and safe imagetext pairs, which we call CoProV2. Using a custom DPO strategy and this dataset, we train safety experts, in the form of low-rank adaptation (LoRA) matrices, able to guide the generation process away from specific safety-related concepts. Then, we merge the experts into a single LoRA using a novel merging strategy for optimal scaling performance. This expert-based approach enables scalability, allowing us to remove 7× more harmful concepts from T2I models compared to baselines. AlignGuard consistently outperforms the state-of-the-art on many benchmarks and establishes new practices for safety alignment in T2I networks. We will release code and models. Warning: this paper includes potentially offensive content.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers9
- Hierarchical Fine-grained Preference Optimization for Physically Plausible Video GenerationHarold Haodong Chen, Haojian Huang, Qifeng Chen, Harry Yang et al.NeurIPS 2025 · 25 citations
- Training-Free Safe Denoisers for Safe Use of Diffusion ModelsMingyu Kim, Dongjun Kim, Amman Yusuf, Stefano Ermon et al.NeurIPS 2025 · 21 citations
- When Safety Collides: Resolving Multi-Category Harmful Conflicts in Text-to-Image Diffusion via Adaptive Safety GuidanceYongli Xiang, Ziming Hong, Zhaoqing Wang, Xiangyu Zhao et al.CVPR 2026 · 14 citations
- DREAM: Scalable Red Teaming for Text-to-Image Generative Systems via Distribution ModelingBoheng Li, Junjie Wang, Yiming Li, Zhiyang Hu et al.S&P 2026 · 9 citations
- What Concepts Lie Within? Detecting and Suppressing Risky Content in Diffusion TransformersChenyu Zhang, Lanjun Wang, Yueyang Cheng, Ruidong Chen et al.CCS 2026 · 1 citation
Builds on28
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann et al.ICLR 2024 · 4,569 citations
Related papers
- Direct Unlearning Optimization for Robust and Safe Text-to-Image ModelsYong-Hyun Park, Sangdoo Yun, Jin-Hwa Kim, Junho Kim et al.NeurIPS 2024 · 60 citations
- ForceForget: Reinforcement Concept Removal for Enhancing Safety in Text-to-Image ModelsDong Han, Yong LiICML 2026
- LoRA-Guard: Parameter-Efficient Guardrail Adaptation for Content Moderation of Large Language ModelsHayder Elesedy, Pedro M. Esperança, Silviu Vlad Oprea, Mete OzayEMNLP 2024 · 5 citations
- GuardT2I: Defending Text-to-Image Models from Adversarial PromptsYijun Yang, Ruiyuan Gao, Xiao Yang, Jianyuan Zhong et al.NeurIPS 2024 · 74 citations
- TarPro: Targeted Protection Against Malicious Image EditingKaixin Shen, Ruijie Quan, Jiaxu Miao, Jun XiaoAAAI 2026
