Learnings from Scaling Visual Tokenizers for Reconstruction and Generation
Philippe Hansen-Estruch, David Yan, Ching-Yao Chuang, Orr Zohar, Jialiang Wang, Tingbo Hou, Tao Xu, Sriram Vishwanath, Peter Vajda, Xinlei Chen
Abstract
Visual tokenization through auto-encoding enhances state-of-the-art image and video generative models by compressing pixels into a latent space. However, questions persist regarding how the design of the auto-encoder affects both reconstruction and downstream generative performance. This paper investigates the impact of scaling autoencoders for reconstruction and generation by substituting the convolutional backbone with an enhanced Vision Transformer for Tokenization (Vi-Tok). This paper's results show that scaling the auto-encoder bottleneck correlates with improved reconstruction, though its relationship with generative performance is more complex. In contrast, scaling the encoder does not lead to gains, while scaling the decoder enhances reconstruction with minimal effect on generation. These findings indicate that scaling the existing autoencoder paradigm does not significantly improve generative performance. When paired with Diffusion Transformers, ViTok achieves competitive image reconstruction & generation performance on 256p and 512p ImageNet-1K. For videos, Vi-Tok achieves state-of-the-art in both reconstruction & generation performance on 128p UCF-101.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2133aa61-e09b-4f14-9e4c-47db59a35135Cited by top-tier papers22
- Diffusion Transformers with Representation AutoencodersBoyang Zheng, Nanye Ma, Shengbang Tong, Saining XieICLR 2026 · 288 citations
- YuE: Scaling Open Foundation Models for Long-Form Music GenerationRuibin Yuan, Hanfeng Lin, Shuyue Guo, Ge Zhang et al.ICLR 2026 · 112 citations
- Aligning Visual Foundation Encoders to Tokenizers for Diffusion ModelsBowei Chen, Sai Bi, Hao Tan, He Zhang et al.ICLR 2026 · 36 citations
- AToken: A Unified Tokenizer for VisionJiasen Lu, Liangchen Song, Mingze Xu, Byeongjoo Ahn et al.CVPR 2026 · 33 citations
- Latent Denoising Makes Good TokenizersJiawei Yang, Tianhong Li, Lijie Fan, Yonglong Tian et al.ICLR 2026 · 17 citations
Builds on27
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
Related papers
- GigaTok: Scaling Visual Tokenizers to 3 Billion Parameters for Autoregressive Image GenerationTianwei Xiong, Jun Hao Liew, Zilong Huang, Jiashi Feng et al.ICCV 2025 · 2 citations
- ViTok-v2: Scaling Native Resolution Autoencoders to 5 Billion ParametersPhilippe Hansen-Estruch, Jiahui Chen, Vivek Ramanujan, Orr Zohar et al.ICML 2026
- Language-Guided Image Tokenization for GenerationKaiwen Zha, Lijun Yu, Alireza Fathi, David A. Ross et al.CVPR 2025
- RecTok: Reconstruction Distillation along Rectified FlowQingyu Shi, Size Wu, Jinbin Bai, Kaidong Yu et al.CVPR 2026 · 5 citations
- Flow to the Mode: Mode-Seeking Diffusion Autoencoders for State-of-the-Art Image TokenizationKyle Sargent, Kyle Hsu, Justin Johnson, Li Fei-Fei et al.ICCV 2025 · 1 citation
