StyleGAN-T: Unlocking the Power of GANs for Fast Large-Scale Text-to-Image Synthesis
Axel Sauer, Tero Karras, Samuli Laine, Andreas Geiger, Timo Aila
Abstract
Text-to-image synthesis has recently seen significant progress thanks to large pretrained language models, large-scale training data, and the introduction of scalable model families such as diffusion and autoregressive models. However, the best-performing models require iterative evaluation to generate a single sample. In contrast, generative adversarial networks (GANs) only need a single forward pass. They are thus much faster, but they currently remain far behind the state-of-the-art in large-scale text-to-image synthesis. This paper aims to identify the necessary steps to regain competitiveness. Our proposed model, StyleGAN-T, addresses the specific requirements of large-scale text-to-image synthesis, such as large capacity, stable training on diverse datasets, strong text alignment, and controllable variation vs. text alignment tradeoff. StyleGAN-T significantly improves over previous GANs and outperforms distilled diffusion models - the previous state-of-the-art in fast text-to-image synthesis - in terms of sample quality and speed.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers86
- Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale PredictionKeyu Tian, Yi Jiang, Zehuan Yuan, Bingyue Peng et al.NeurIPS 2024 · 1,199 citations
- Improved Distribution Matching Distillation for Fast Image SynthesisTianwei Yin, Michaël Gharbi, Taesung Park, Richard Zhang et al.NeurIPS 2024 · 728 citations
- ELITE: Encoding Visual Concepts into Textual Embeddings for Customized Text-to-Image GenerationYuxiang Wei, Yabo Zhang, Zhilong Ji, Jinfeng Bai et al.ICCV 2023 · 469 citations
- ControlVideo: Training-free Controllable Text-to-video GenerationYabo Zhang, Yuxiang Wei, Dongsheng Jiang, Xiaopeng Zhang et al.ICLR 2024 · 359 citations
- InstaFlow: One Step is Enough for High-Quality Diffusion-Based Text-to-Image GenerationXingchao Liu, Xiwen Zhang, Jianzhu Ma, Jian Peng et al.ICLR 2024 · 358 citations
Builds on27
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
Related papers
- Scaling up GANs for Text-to-Image SynthesisMinguk Kang, Jun-Yan Zhu, Richard Zhang, Jaesik Park et al.CVPR 2023
- Style Aligned Image Generation via Shared AttentionAmir Hertz, Andrey Voynov, Shlomi Fruchter, Daniel Cohen-OrCVPR 2024
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- UFOGen: You Forward Once Large Scale Text-to-Image Generation via Diffusion GANsYanwu Xu, Yang Zhao, Zhisheng Xiao, Tingbo HouCVPR 2024
- SwiftBrush: One-Step Text-to-Image Diffusion Model with Variational Score DistillationThuan Hoang Nguyen, Anh TranCVPR 2024 · 20 citations
