Würstchen: An Efficient Architecture for Large-Scale Text-to-Image Diffusion Models
Pablo Pernias, Dominic Rampas, Mats Leon Richter, Christopher Pal, Marc Aubreville
Abstract
We introduce Würstchen, a novel architecture for text-to-image synthesis that combines competitive performance with unprecedented cost-effectiveness for largescale text-to-image diffusion models. A key contribution of our work is to develop a latent diffusion technique in which we learn a detailed but extremely compact semantic image representation used to guide the diffusion process. This highly compressed representation of an image provides much more detailed guidance compared to latent representations of language and this significantly reduces the computational requirements to achieve state-of-the-art results. Our approach also improves the quality of text-conditioned image generation based on our user preference study. The training requirements of our approach consists of 24,602 A100-GPU hours -compared to Stable Diffusion 2.1's 200,000 GPU hours. Our approach also requires less training data to achieve these results. Furthermore, our compact latent representations allows us to perform inference over twice as fast, slashing the usual costs and carbon footprint of a state-of-the-art (SOTA) diffusion model significantly, without compromising the end performance. In a broader comparison against SOTA models our approach is substantially more efficient and compares favourably in terms of image quality. We believe that this work motivates more emphasis on the prioritization of both performance and computational accessibility. Figure 1: Text-conditional generations using Würstchen. Note the various art styles and aspect ratios.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4b9979cd-198b-4a19-9136-396c52c586cfCited by top-tier papers59
- What matters for Representation Alignment: Global Information or Spatial Structure?Jaskirat Singh, Xingjian Leng, Zongze Wu, Liang Zheng et al.ICLR 2026 · 84 citations
- PixelDiT: Pixel Diffusion Transformers for Image GenerationYongsheng Yu, Wei Xiong, Weili Nie, Yichen Sheng et al.CVPR 2026 · 82 citations
- UltraPixel: Advancing Ultra High-Resolution Image Synthesis to New PeaksJingjing Ren, Wenbo Li, Haoyu Chen, Renjing Pei et al.NeurIPS 2024 · 74 citations
- Direct Unlearning Optimization for Robust and Safe Text-to-Image ModelsYong-Hyun Park, Sangdoo Yun, Jin-Hwa Kim, Junho Kim et al.NeurIPS 2024 · 60 citations
- Token Merging for Training-Free Semantic Binding in Text-to-Image SynthesisTaihang Hu, Linxuan Li, Joost van de Weijer, Hongcheng Gao et al.NeurIPS 2024 · 45 citations
Builds on19
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
Related papers
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann et al.ICLR 2024 · 4,569 citations
- SnapFusion: Text-to-Image Diffusion Model on Mobile Devices within Two SecondsYanyu Li, Huan Wang, Qing Jin, Ju Hu et al.NeurIPS 2023 · 300 citations
- On improved Conditioning Mechanisms and Pre-training Strategies for Diffusion ModelsTariq Berrada Ifriqi, Pietro Astolfi, Melissa Hall, Reyhane Askari Hemmat et al.NeurIPS 2024 · 18 citations
- PaGoDA: Progressive Growing of a One-Step Generator from a Low-Resolution Diffusion TeacherDongjun Kim, Chieh-Hsin Lai, Wei-Hsiang Liao, Yuhta Takida et al.NeurIPS 2024 · 32 citations
- KOALA: Empirical Lessons Toward Memory-Efficient and Fast Diffusion Models for Text-to-Image SynthesisYoungwan Lee, Kwanyong Park, Yoorhim Cho, Yong-Ju Lee et al.NeurIPS 2024 · 22 citations
