One-Step is Enough: Sparse Autoencoders for Text-to-Image Diffusion Models
Viacheslav Surkov, Chris Wendler, Antonio Mari, Mikhail Terekhov, Justin Deschenaux, Robert West, Caglar Gulcehre, David Bau
Abstract
For large language models (LLMs), sparse autoencoders (SAEs) have been shown to decompose intermediate representations that often are not interpretable directly into sparse sums of interpretable features, facilitating better control and subsequent analysis. However, similar analyses and approaches have been lacking for text-toimage models. We investigate the possibility of using SAEs to learn interpretable features for SDXL Turbo, a few-step text-to-image diffusion model. To this end, we train SAEs on the updates performed by transformer blocks within SDXL Turbo's denoising U-net in its 1-step setting. Interestingly, we find that they generalize to 4-step SDXL Turbo and even to the multi-step SDXL base model (i.e., a different model) without additional training. In addition, we show that their learned features are interpretable, causally influence the generation process, and reveal specialization among the blocks. We do so by creating RIEBench, a representation-based image editing benchmark, for editing images while they are generated by turning on and off individual SAE features. This allows us to track which transformer blocks' features are the most impactful depending on the edit category. Our work is the first investigation of SAEs for interpretability in text-toimage diffusion models and our results establish SAEs as a promising approach for understanding and manipulating the internal mechanisms of text-to-image models.
39th Conference on Neural Information Processing Systems (NeurIPS 2025).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e7a2137a-2d72-440f-a3c3-dc7bcfe9fa2dCited by top-tier papers14
- Sparse Autoencoders Learn Monosemantic Features in Vision-Language ModelsMateusz Pach, Shyamgopal Karthik, Quentin Bouniot, Serge J. Belongie et al.NeurIPS 2025 · 79 citations
- Projecting Assumptions: The Duality Between Sparse Autoencoders and Concept GeometrySai Sumedh R. Hindupur, Ekdeep Singh Lubana, Thomas Fel, Demba BaNeurIPS 2025 · 65 citations
- When Are Concepts Erased From Diffusion Models?Kevin Lu, Nicky Kriplani, Rohit Gandikota, Minh Pham et al.NeurIPS 2025 · 21 citations
- Interpretable Debiasing of Vision-Language Models for Social FairnessNa Min An, Yoonna Jang, Yusuke Hirota, Ryo Hachiuma et al.CVPR 2026 · 7 citations
- SAEmnesia: Erasing Concepts in Diffusion Models with Supervised Sparse AutoencodersEnrico Cassano, Riccardo Renzulli, Marco Nurisso, Mirko Zaffaroni et al.ICML 2026 · 7 citations
Builds on27
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
Related papers
- Emergence and Evolution of Interpretable Concepts in Diffusion ModelsBerk Tinaz, Zalan Fabian, Mahdi SoltanolkotabiNeurIPS 2025 · 23 citations
- DLM-Scope: Mechanistic Interpretability of Diffusion Language Models via Sparse AutoencodersXu Wang, Bingqing Jiang, Yu Wan, Baosong Yang et al.ICML 2026
- Revelio: Interpreting and Leveraging Semantic Information in Diffusion ModelsDahye Kim, Xavier Thomas, Deepti GhadiyaramICCV 2025 · 1 citation
- TIDE: Temporal-Aware Sparse Autoencoders for Interpretable Diffusion Transformers in Image GenerationVictor Shea-Jay Huang, Le Zhuo, Yi Xin, Zhaokai Wang et al.AAAI 2026 · 10 citations
- Discovering and Steering Interpretable Concepts in Large Generative Music ModelsNikhil Singh, Manuel Cherep, Pattie MaesICLR 2026 · 17 citations
