Self-Discovering Interpretable Diffusion Latent Directions for Responsible Text-to-Image Generation
Hang Li, Chengzhi Shen, Philip Torr, Volker Tresp, Jindong Gu
Abstract
Diffusion-based models have gained significant popularity for text-to-image generation due to their exceptional image-generation capabilities. A risk with these models is the potential generation of inappropriate content, such as biased or harmful images. However, the underlying reasons for generating such undesired content from the perspective of the diffusion model's internal representation remain unclear. Previous work interprets vectors in an interpretable latent space of diffusion models as semantic concepts. However, existing approaches cannot discover directions for arbitrary concepts, such as those related to inappropriate concepts. In this work, we propose a novel self-supervised approach to find interpretable latent directions for a given concept. With the discovered vectors, we further propose a simple approach to mitigate inappropriate generation. Extensive experiments have been conducted to verify the effectiveness of our mitigation approach, namely, for fair generation, safe generation, and responsible text-enhancing generation. Project page: https://interpretdiffusion.github.io.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bf2c321b-a691-4ce9-9fd0-50f9b5c94cd5Cited by top-tier papers39
- One-Step is Enough: Sparse Autoencoders for Text-to-Image Diffusion ModelsViacheslav Surkov, Chris Wendler, Antonio Mari, Mikhail Terekhov et al.NeurIPS 2025 · 33 citations
- CURE: Concept Unlearning via Orthogonal Representation Editing in Diffusion ModelsShristi Das Biswas, Arani Roy, Kaushik RoyNeurIPS 2025 · 30 citations
- How to Trace Latent Generative Model Generated Images without Artificial Watermark?Zhenting Wang, Vikash Sehwag, Chen Chen, Lingjuan Lyu et al.ICML 2024 · 24 citations
- Diffusion PID: Interpreting Diffusion via Partial Information DecompositionShaurya Dewan, Rushikesh Zawar, Prakanshul Saxena, Yingshan Chang et al.NeurIPS 2024 · 23 citations
- Incomplete Multi-view Clustering via Diffusion Contrastive GenerationYuanyang Zhang, Yijie Lin, Weiqing Yan, Li Yao et al.AAAI 2025 · 19 citations
Builds on28
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
Related papers
- Responsible Diffusion Models via Constraining Text Embeddings within Safe RegionsZhiwen Li, Die Chen, Mingyuan Fan, Cen Chen et al.WWW 2025 · 11 citations
- Plug-and-Play Interpretable Responsible Text-to-Image Generation via Dual-Space Multi-facet Concept ControlBasim Azam, Naveed AkhtarCVPR 2025
- Responsible Text-to-Image Diffusion: Interpretable and Linearly Controllable Semantics for Fair and Safe GenerationSayedmoslem Shokrolahi, Jae-Mo Kang, Il-Min KimICML 2026
- Detect-and-Guide: Self-regulation of Diffusion Models for Safe Text-to-Image Generation via Guideline Token OptimizationFeifei Li, Mi Zhang, Yiming Sun, Min YangCVPR 2025
- Training-Free Safe Text Embedding Guidance for Text-to-Image Diffusion ModelsByeonghu Na, Mina Kang, Jiseok Kwak, Minsang Park et al.NeurIPS 2025 · 8 citations
