ViTGAN: Training GANs with Vision Transformers
Kwonjoon Lee, Huiwen Chang, Lu Jiang, Han Zhang, Zhuowen Tu, Ce Liu
Abstract
Recently, Vision Transformers (ViTs) have shown competitive performance on image recognition while requiring less vision-specific inductive biases. In this paper, we investigate if such performance can be extended to image generation. To this end, we integrate the ViT architecture into generative adversarial networks (GANs). For ViT discriminators, we observe that existing regularization methods for GANs interact poorly with self-attention, causing serious instability during training. To resolve this issue, we introduce several novel regularization techniques for training GANs with ViTs. For ViT generators, we examine architectural choices for latent and pixel mapping layers to facilitate convergence. Empirically, our approach, named ViTGAN, achieves comparable performance to the leading CNNbased GAN models on three datasets: CIFAR-10, CelebA, and LSUN bedroom. Our code is available online 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5a552f8a-9160-4da7-9126-624f7b3ee7beCited by top-tier papers23
- StyleSwin: Transformer-based GAN for High-resolution Image GenerationBowen Zhang, Shuyang Gu, Bo Zhang, Jianmin Bao et al.CVPR 2022 · 217 citations
- Poisson Flow Generative ModelsYilun Xu, Ziming Liu, Max Tegmark, Tommi S. JaakkolaNeurIPS 2022 · 133 citations
- ActFormer: A GAN-based Transformer towards General Action-Conditioned 3D Human Motion GenerationLiang Xu, Ziyang Song, Dongliang Wang, Jing Su et al.ICCV 2023 · 100 citations
- Reflected Diffusion ModelsAaron Lou, Stefano ErmonICML 2023 · 84 citations
- One-Step Diffusion with Distribution Matching DistillationTianwei Yin, Michaël Gharbi, Richard Zhang, Eli Shechtman et al.CVPR 2024 · 75 citations
Builds on32
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
Related papers
- TransGAN: Two Pure Transformers Can Make One Strong GAN, and That Can Scale UpYifan Jiang, Shiyu Chang, Zhangyang WangNeurIPS 2021 · 515 citations
- All are Worth Words: A ViT Backbone for Diffusion ModelsFan Bao, Shen Nie, Kaiwen Xue, Yue Cao et al.CVPR 2023
- When Adversarial Training Meets Vision Transformers: Recipes from Training to ArchitectureYichuan Mo, Dongxian Wu, Yifei Wang, Yiwen Guo et al.NeurIPS 2022 · 109 citations
- Intriguing Properties of Vision TransformersMuzammal Naseer, Kanchana Ranasinghe, Salman Khan, Munawar Hayat et al.NeurIPS 2021 · 863 citations
- Are Transformers more robust than CNNs?Yutong Bai, Jieru Mei, Alan L. Yuille, Cihang XieNeurIPS 2021 · 365 citations
