Improved Transformer for High-Resolution GANs
Long Zhao, Zizhao Zhang, Ting Chen, Dimitris N. Metaxas, Han Zhang
摘要
Attention-based models, exemplified by the Transformer, can effectively model long range dependency, but suffer from the quadratic complexity of self-attention operation, making them difficult to be adopted for high-resolution image generation based on Generative Adversarial Networks (GANs). In this paper, we introduce two key ingredients to Transformer to address this challenge. First, in low-resolution stages of the generative process, standard global self-attention is replaced with the proposed multi-axis blocked self-attention which allows efficient mixing of local and global attention. Second, in high-resolution stages, we drop self-attention while only keeping multi-layer perceptrons reminiscent of the implicit neural function. To further improve the performance, we introduce an additional self-modulation component based on cross-attention. The resulting model, denoted as HiT, has a nearly linear computational complexity with respect to the image size and thus directly scales to synthesizing high definition images. We show in the experiments that the proposed HiT achieves state-of-the-art FID scores of 30.83 and 2.95 on unconditional ImageNet and FFHQ , respectively, with a reasonable throughput. We believe the proposed HiT is an important milestone for generators in GANs which are completely free of convolutions. Our code is made publicly available at https://github.com/google-research/hit-gan
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper19
- Uformer: A General U-Shaped Transformer for Image RestorationZhendong Wang, Xiaodong Cun, Jianmin Bao, Wengang Zhou 等CVPR 2022 · 被引用 1,970 次
- Muse: Text-To-Image Generation via Masked Generative TransformersHuiwen Chang, Han Zhang, Jarred Barber, Aaron Maschinot 等ICML 2023 · 被引用 751 次
- MAXIM: Multi-Axis MLP for Image ProcessingZhengzhong Tu, Hossein Talebi, Han Zhang, Feng Yang 等CVPR 2022 · 被引用 550 次
- TransGAN: Two Pure Transformers Can Make One Strong GAN, and That Can Scale UpYifan Jiang, Shiyu Chang, Zhangyang WangNeurIPS 2021 · 被引用 515 次
- StyleSwin: Transformer-based GAN for High-resolution Image GenerationBowen Zhang, Shuyang Gu, Bo Zhang, Jianmin Bao 等CVPR 2022 · 被引用 217 次
它引用的顶会 Paper24
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray 等ICML 2021 · 被引用 6,356 次
相关 Paper
- Swin-UNIT: Transformer-based GAN for High-resolution Unpaired Image TranslationYifan Li, Yaochen Li, Wenneng Tang, Zhifeng Zhu 等ACM MM 2023 · 被引用 13 次
- Holistic Tokenizer for Autoregressive Image GenerationAnlin Zheng, Haochen Wang, Yucheng Zhao, Weipeng Deng 等ICCV 2025 · 被引用 11 次
- Styleformer: Transformer based Generative Adversarial Networks with Style VectorJeeseung Park, Younggeun KimCVPR 2022 · 被引用 49 次
- Recursive Generalization Transformer for Image Super-ResolutionZheng Chen, Yulun Zhang, Jinjin Gu, Linghe Kong 等ICLR 2024 · 被引用 81 次
- Scalable High-Resolution Pixel-Space Image Synthesis with Hourglass Diffusion TransformersKatherine Crowson, Stefan Andreas Baumann, Alex Birch, Tanishq Mathew Abraham 等ICML 2024 · 被引用 98 次
