Improving Visual Quality of Image Synthesis by A Token-based Generator with Transformers
Yanhong Zeng, Huan Yang, Hongyang Chao, Jianbo Wang, Jianlong Fu
Abstract
We present a new perspective of achieving image synthesis by viewing this task as a visual token generation problem. Different from existing paradigms that directly synthesize a full image from a single input (e.g., a latent code), the new formulation enables a flexible local manipulation for different image regions, which makes it possible to learn content-aware and fine-grained style control for image synthesis. Specifically, it takes as input a sequence of latent tokens to predict the visual tokens for synthesizing an image. Under this perspective, we propose a token-based generator (i.e.,TokenGAN). Particularly, the TokenGAN inputs two semantically different visual tokens, i.e., the learned constant content tokens and the style tokens from the latent space. Given a sequence of style tokens, the TokenGAN is able to control the image synthesis by assigning the styles to the content tokens by attention mechanism with a Transformer. We conduct extensive experiments and show that the proposed TokenGAN has achieved state-of-the-art results on several widelyused image synthesis benchmarks, including FFHQ and LSUN CHURCH with different resolutions. In particular, the generator is able to synthesize high-fidelity images with 1024 × 1024 size, dispensing with convolutions entirely.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers10
- EfficientFormer: Vision Transformers at MobileNet SpeedYanyu Li, Geng Yuan, Yang Wen, Ju Hu et al.NeurIPS 2022 · 742 citations
- Rethinking Vision Transformers for MobileNet Size and SpeedYanyu Li, Ju Hu, Yang Wen, Georgios Evangelidis et al.ICCV 2023 · 300 citations
- Class-Aware Adversarial Transformers for Medical Image SegmentationChenyu You, Ruihan Zhao, Fenglin Liu, Siyuan Dong et al.NeurIPS 2022 · 137 citations
- Learning Trajectory-Aware Transformer for Video Super-ResolutionChengxu Liu, Huan Yang, Jianlong Fu, Xueming QianCVPR 2022 · 113 citations
- Advancing High-Resolution Video-Language Representation with Large-Scale Video TranscriptionsHongwei Xue, Tiankai Hang, Yanhong Zeng, Yuchong Sun et al.CVPR 2022 · 107 citations
Builds on25
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- Training Generative Adversarial Networks with Limited DataTero Karras, Miika Aittala, Janne Hellsten, Samuli Laine et al.NeurIPS 2020 · 2,345 citations
Related papers
- Holistic Tokenizer for Autoregressive Image GenerationAnlin Zheng, Haochen Wang, Yucheng Zhao, Weipeng Deng et al.ICCV 2025 · 11 citations
- SemanticStyleGAN: Learning Compositional Generative Priors for Controllable Image Synthesis and EditingYichun Shi, Xiao Yang, Yangyue Wan, Xiaohui ShenCVPR 2022 · 88 citations
- Autoregression with Self-Token PredictionDengsheng Chen, Yangming Shi, Enhua WuICML 2026
- FOCUS: Unified Vision-Language Modeling for Interactive Editing Driven by Referential SegmentationFan Yang, Yousong Zhu, Xin Li, Yufei Zhan et al.NeurIPS 2025 · 6 citations
- VibeToken: Scaling 1D Image Tokenizers and Autoregressive Models for Dynamic Resolution GenerationsMaitreya Patel, Jingtao Li, Weiming Zhuang, Yezhou Yang et al.CVPR 2026 · 2 citations
