Masked Autoencoders Are Effective Tokenizers for Diffusion Models
Hao Chen, Yujin Han, Fangyi Chen, Xiang Li, Yidong Wang, Jindong Wang, Ze Wang, Zicheng Liu, Difan Zou, Bhiksha Raj
Abstract
Recent advances in latent diffusion models have demonstrated their effectiveness for highresolution image synthesis. However, the properties of the latent space from tokenizer for better learning and generation of diffusion models remain under-explored. Theoretically and empirically, we find that improved generation quality is closely tied to the latent distributions with better structure, such as the ones with fewer Gaussian Mixture modes and more discriminative features. Motivated by these insights, we propose MAE-Tok, an autoencoder (AE) leveraging mask modeling to learn semantically rich latent space while maintaining reconstruction fidelity. Extensive experiments validate our analysis, demonstrating that the variational form of autoencoders is not necessary, and a discriminative latent space from AE alone enables state-of-the-art performance on ImageNet generation using only 128 tokens. MAETok achieves significant practical improvements, enabling a gFID of 1.69 with 76× faster training and 31× higher inference throughput for 512×512 generation. Our findings show that the structure of the latent space, rather than variational constraints, is crucial for effective diffusion models. Code and trained models are released 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1648d4fa-c9d2-4e70-8853-b4eb8219f0dfCited by top-tier papers34
- Representation Alignment for Diffusion Transformers without External ComponentsDengyang Jiang, Mengmeng Wang, Liuzhuozheng Li, Lei Zhang et al.ICLR 2026 · 532 citations
- Diffusion Transformers with Representation AutoencodersBoyang Zheng, Nanye Ma, Shengbang Tong, Saining XieICLR 2026 · 288 citations
- UniTok: a Unified Tokenizer for Visual Generation and UnderstandingChuofan Ma, Yi Jiang, Junfeng Wu, Jihan Yang et al.NeurIPS 2025 · 164 citations
- PixelDiT: Pixel Diffusion Transformers for Image GenerationYongsheng Yu, Wei Xiong, Weili Nie, Yichen Sheng et al.CVPR 2026 · 82 citations
- Aligning Visual Foundation Encoders to Tokenizers for Diffusion ModelsBowei Chen, Sai Bi, Hao Tan, He Zhang et al.ICLR 2026 · 36 citations
Builds on45
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
Related papers
- Latent Denoising Makes Good TokenizersJiawei Yang, Tianhong Li, Lijie Fan, Yonglong Tian et al.ICLR 2026 · 17 citations
- VideoMAETok: Boosting Video Diffusion Models via Masked Autoencoders as TokenizersZhan Tong, Tinne TuytelaarsICML 2026
- Towards Sequence Modeling Alignment between Tokenizer and Autoregressive ModelPingyu Wu, Kai Zhu, Yu Liu, Longxiang Tang et al.ICLR 2026 · 16 citations
- Stabilize the Latent Space for Image Autoregressive Modeling: A Unified PerspectiveYongxin Zhu, Bocheng Li, Hang Zhang, Xin Li et al.NeurIPS 2024 · 26 citations
- Unified Latent Space for Understanding and Generation via Semantic Auto-encoderXiaojie Li, Yang Zhao, Ming Li, Yancheng Zhang et al.CVPR 2026
