Latent Diffusion Models With Masked Autoencoders
Junho Lee, Jeongwoo Shin, Hyungwook Choi, Joonseok Lee
Abstract
In spite of the remarkable potential of Latent Diffusion Models (LDMs) in image generation, the desired properties and optimal design of the autoencoders have been underexplored. In this work, we analyze the role of autoencoders in LDMs and identify three key properties: latent smoothness, perceptual compression quality, and reconstruction quality. We demonstrate that existing autoencoders fail to simultaneously satisfy all three properties, and propose Variational Masked AutoEncoders (VMAEs), taking advantage of the hierarchical features maintained by Masked AutoEncoders. We integrate VMAEs into the LDM framework, introducing Latent Diffusion Models with Masked AutoEncoders (LDMAEs). Through comprehensive experiments, we demonstrate significantly enhanced image generation quality and computational efficiency. Our code is available at https://github.com/isno0907/ldmae.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2627664e-7dfc-4e57-bf79-5e957c09b297Cited by top-tier papers6
- Aligning Visual Foundation Encoders to Tokenizers for Diffusion ModelsBowei Chen, Sai Bi, Hao Tan, He Zhang et al.ICLR 2026 · 36 citations
- Toward Diffusible High-Dimensional Latent Spaces: A Frequency PerspectiveBolin Lai, Xudong Wang, Saketh Rambhatla, James M. Rehg et al.CVPR 2026 · 7 citations
- Geometry-Aware Image Flow MatchingJunho Lee, Kwanseok Kim, Joonseok LeeICML 2026 · 3 citations
- MoLingo: Motion-Language Alignment for Text-to-Human Motion GenerationYannan He, Garvita Tiwari, Xiaohan Zhang, Pankaj Bora et al.CVPR 2026 · 2 citations
- Distribution Matching Variational AutoEncoderSen Ye, Jianning Pei, Mengde Xu, Shuyang Gu et al.ICML 2026
Builds on26
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
Related papers
- LMD: Faster Image Reconstruction with Latent Masking DiffusionZhiyuan Ma, Zhihuan Yu, Jianjun Li, Bowen ZhouAAAI 2024 · 15 citations
- Variational Autoencoding Discrete Diffusion with Enhanced Dimensional Correlations ModelingTianyu Xie, Shuchen Xue, Zijin Feng, Tianyang Hu et al.ICLR 2026 · 14 citations
- LiteVAE: Lightweight and Efficient Variational Autoencoders for Latent Diffusion ModelsSeyedmorteza Sadat, Jakob Buhmann, Derek Bradley, Otmar Hilliges et al.NeurIPS 2024 · 29 citations
- Boosting Latent Diffusion with Perceptual ObjectivesTariq Berrada, Pietro Astolfi, Melissa Hall, Marton Havasi et al.ICLR 2025
- Optimal Stopping in Latent Diffusion ModelsYu-Han Wu, Quentin Berthet, Gérard Biau, Claire Boyer et al.ICML 2026 · 1 citation
