Lune

ICML2025Top-tier venue

Learnings from Scaling Visual Tokenizers for Reconstruction and Generation

Philippe Hansen-Estruch, David Yan, Ching-Yao Chuang, Orr Zohar, Jialiang Wang, Tingbo Hou, Tao Xu, Sriram Vishwanath, Peter Vajda, Xinlei Chen

2025Year
22Top-tier citations

Abstract

Visual tokenization through auto-encoding enhances state-of-the-art image and video generative models by compressing pixels into a latent space. However, questions persist regarding how the design of the auto-encoder affects both reconstruction and downstream generative performance. This paper investigates the impact of scaling autoencoders for reconstruction and generation by substituting the convolutional backbone with an enhanced Vision Transformer for Tokenization (Vi-Tok). This paper's results show that scaling the auto-encoder bottleneck correlates with improved reconstruction, though its relationship with generative performance is more complex. In contrast, scaling the encoder does not lead to gains, while scaling the decoder enhances reconstruction with minimal effect on generation. These findings indicate that scaling the existing autoencoder paradigm does not significantly improve generative performance. When paired with Diffusion Transformers, ViTok achieves competitive image reconstruction & generation performance on 256p and 512p ImageNet-1K. For videos, Vi-Tok achieves state-of-the-art in both reconstruction & generation performance on 128p UCF-101.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 2133aa61-e09b-4f14-9e4c-47db59a35135

Cited by top-tier papers22

Ask how each one uses it

Builds on27

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines