GigaTok: Scaling Visual Tokenizers to 3 Billion Parameters for Autoregressive Image Generation
Tianwei Xiong, Jun Hao Liew, Zilong Huang, Jiashi Feng, Xihui Liu
Abstract
In autoregressive (AR) image generation, visual tokenizers compress images into compact discrete latent tokens, enabling efficient training of downstream autoregressive models for visual generation via next-token prediction. While scaling visual tokenizers improves image reconstruction quality, it often degrades downstream generation qual-ity-a challenge not adequately addressed in existing literature. To address this, we introduce GigaTok, the first approach to simultaneously improve image reconstruction, generation, and representation learning when scaling visual tokenizers. We identify the growing complexity of latent space as the key factor behind the reconstruction vs. generation dilemma. To mitigate this, we propose semantic regularization, which aligns tokenizer features with semantically consistent features from a pre-trained visual encoder. This constraint prevents excessive latent space complexity during scaling, yielding consistent improvements in both reconstruction and downstream autoregressive generation. Building on semantic regularization, we explore three key practices for scaling tokenizers: (1) using 1D tokenizers for better scalability, (2) prioritizing decoder scaling when expanding both encoder and decoder, and (3) employing entropy loss to stabilize training for billion-scale tokenizers. By scaling to billion parameters, GigaTok achieves state-of-the-art performance in reconstruction, downstream generation, and downstream AR representation quality.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 190b22e2-ca67-4b23-93fd-ae38b714eb71Cited by top-tier papers19
- Aligning Visual Foundation Encoders to Tokenizers for Diffusion ModelsBowei Chen, Sai Bi, Hao Tan, He Zhang et al.ICLR 2026 · 36 citations
- AToken: A Unified Tokenizer for VisionJiasen Lu, Liangchen Song, Mingze Xu, Byeongjoo Ahn et al.CVPR 2026 · 33 citations
- Towards Sequence Modeling Alignment between Tokenizer and Autoregressive ModelPingyu Wu, Kai Zhu, Yu Liu, Longxiang Tang et al.ICLR 2026 · 16 citations
- WeTok: Powerful Discrete Tokenization for High-Fidelity Visual ReconstructionShaobin Zhuang, Yiwei Guo, Fangyikang Wang, Canmiao Fu et al.ICLR 2026 · 9 citations
- EVATok: Adaptive Length Video Tokenization for Efficient Visual Autoregressive GenerationTianwei Xiong, Jun Hao Liew, Zilong Huang, Zhijie Lin et al.CVPR 2026 · 8 citations
Builds on43
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
Related papers
- Learnings from Scaling Visual Tokenizers for Reconstruction and GenerationPhilippe Hansen-Estruch, David Yan, Ching-Yao Chuang, Orr Zohar et al.ICML 2025
- ViTok-v2: Scaling Native Resolution Autoencoders to 5 Billion ParametersPhilippe Hansen-Estruch, Jiahui Chen, Vivek Ramanujan, Orr Zohar et al.ICML 2026
- ImageFolder: Autoregressive Image Generation with Folded TokensXiang Li, Kai Qiu, Hao Chen, Jason Kuen et al.ICLR 2025
- When Worse is Better: Navigating the Compression Generation Trade-off In Visual TokenizationVivek Ramanujan, Kushal Tirumala, Armen Aghajanyan, Luke Zettlemoyer et al.NeurIPS 2025 · 3 citations
- Prompt Yourself: Awakening Textual Semantics in 1D Visual TokenizersHualiang Wang, Siming Fu, Weinan Jia, Yuning Lu et al.CVPR 2026
