LMD: Faster Image Reconstruction with Latent Masking Diffusion
Zhiyuan Ma, Zhihuan Yu, Jianjun Li, Bowen Zhou
Abstract
As a class of fruitful approaches, diffusion probabilistic models (DPMs) have shown excellent advantages in highresolution image reconstruction. On the other hand, masked autoencoders (MAEs), as popular self-supervised vision learners, have demonstrated simpler and more effective image reconstruction and transfer capabilities on downstream tasks. However, they all require extremely high training costs, either due to inherent high temporal-dependence (i.e., excessively long diffusion steps) or due to artificially low spatialdependence (i.e., human-formulated high mask ratio, such as 0.75). To the end, this paper presents LMD, a simple but faster image reconstruction framework with Latent Masking Diffusion. First, we propose to project and reconstruct images in latent space through a pre-trained variational autoencoder, which is theoretically more efficient than in the pixel-based space. Then, we combine the advantages of MAEs and DPMs to design a progressive masking diffusion model, which gradually increases the masking proportion by three different schedulers and reconstructs the latent features from simple to difficult, without sequentially performing denoising diffusion as in DPMs or using fixed high masking ratio as in MAEs, so as to alleviate the high training time-consumption predicament. Our approach allows for learning high-capacity models and accelerate their training (by 3× or more) and barely reduces the original accuracy. Inference speed in downstream tasks also significantly outperforms the previous approaches 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f9afb241-72b7-40f7-87a5-ff7ea1b390c5Cited by top-tier papers7
- Follow Your Pose: Pose-Guided Text-to-Video Generation Using Pose-Free VideosYue Ma, Yingqing He, Xiaodong Cun, Xintao Wang et al.AAAI 2024 · 318 citations
- Neural Residual Diffusion Models for Deep Scalable Vision GenerationZhiyuan Ma, Liangliang Zhao, Biqing Qi, Bowen ZhouNeurIPS 2024 · 15 citations
- Safe-SD: Safe and Traceable Stable Diffusion with Text Prompt Trigger for Invisible Generative WatermarkingZhiyuan Ma, Guoli Jia, Biqing Qi, Bowen ZhouACM MM 2024 · 14 citations
- AdapEdit: Spatio-Temporal Guided Adaptive Editing Algorithm for Text-Based Continuity-Sensitive Image EditingZhiyuan Ma, Guoli Jia, Bowen ZhouAAAI 2024 · 13 citations
- MoleBridge: Synthetic Space Projecting with Discrete Markov BridgesRongchao Zhang, Yu Huang, Yongzhi Cao, Hanpin WangNeurIPS 2025 · 8 citations
Builds on23
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
Related papers
- Latent Diffusion Models With Masked AutoencodersJunho Lee, Jeongwoo Shin, Hyungwook Choi, Joonseok LeeICCV 2025 · 3 citations
- Progressively Compressed Auto-Encoder for Self-supervised Representation LearningJin Li, Yaoming Wang, Xiaopeng Zhang, Yabo Chen et al.ICLR 2023
- Understanding Masked Autoencoders via Hierarchical Latent Variable ModelsLingjing Kong, Martin Q. Ma, Guangyi Chen, Eric P. Xing et al.CVPR 2023
- Masked Autoencoders Are Scalable Vision LearnersKaiming He, Xinlei Chen, Saining Xie, Yanghao Li et al.CVPR 2022
- USP: Unified Self-Supervised Pretraining for Image Generation and UnderstandingXiangxiang Chu, Renda Li, Yong WangICCV 2025 · 3 citations
