Improving Training Efficiency of Diffusion Models via Multi-Stage Framework and Tailored Multi-Decoder Architecture
Huijie Zhang, Yifu Lu, Ismail Alkhouri, Saiprasad Ravishankar, Dogyoon Song, Qing Qu
Abstract
Diffusion models, emerging as powerful deep generative tools, excel in various applications. They operate through a two-steps process: introducing noise into training samples and then employing a model to convert random noise into new samples (e.g., images). However, their remarkable generative performance is hindered by slow training and sampling. This is due to the necessity of tracking extensive forward and reverse diffusion trajectories, and employing a large model with numerous parameters across multiple timesteps (i.e., noise levels). To tackle these challenges, we present a multi-stage framework inspired by our empirical findings. These observations indicate the advantages of employing distinct parameters tailored to each timestep while retaining universal parameters shared across all time steps. Our approach involves segmenting the time interval into multiple stages where we employ custom multi-decoder U-net architecture that blends time-dependent models with a universally shared encoder. Our framework enables the efficient distribution of computational resources and mitigates inter-stage interference, which substantially improves training efficiency. Extensive numerical experiments affirm the effectiveness of our framework, showcasing significant training and sampling efficiency enhancements on three state-of-the-art diffusion models, including large-scale latent diffusion models. Furthermore, our ablation studies illustrate the impact of two important components in our framework: (i) a novel timestep clustering algorithm for stage division, and (ii) an innovative multi-decoder U-net architecture, seamlessly integrating universal and customized hyperparameters.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d9775192-fe7a-48bc-9a00-5c520a3b3a45Cited by top-tier papers7
- DepthFM: Fast Generative Monocular Depth Estimation with Flow MatchingMing Gui, Johannes Schusterbauer, Ulrich Prestel, Pingchuan Ma et al.AAAI 2025 · 51 citations
- UGoDIT: Unsupervised Group Deep Image Prior Via Transferable WeightsShijun Liang, Ismail Alkhouri, Siddhant Gautam, Qing Qu et al.NeurIPS 2025 · 3 citations
- Efficient Training for Human Video Generation with Entropy-Guided Prioritized Progressive LearningChanglin Li, Jiawei Zhang, Shuhao Liu, Sihao Lin et al.CVPR 2026 · 2 citations
- Joint Learning of Pose Regression and Denoising Diffusion with Score Scaling Sampling for Category-Level 6D Pose EstimationSeunghyun Lee, Tae-Kyun KimICCV 2025 · 2 citations
- TimeStep Master: Asymmetrical Mixture of Timestep LoRA Experts for Versatile and Efficient Diffusion Models in VisionShaobin Zhuang, Yiwei Guo, Yanbo Ding, Kunchang Li et al.ICML 2025
Builds on25
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
- GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion ModelsAlexander Quinn Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam et al.ICML 2022 · 4,691 citations
Related papers
- Decouple-Then-Merge: Finetune Diffusion Models as Multi-Task LearningQianli Ma, Xuefei Ning, Dongrui Liu, Li Niu et al.CVPR 2025
- TPDiff: Temporal Pyramid Video Diffusion ModelLingmin Ran, Mike Zheng ShouICLR 2026 · 4 citations
- Nested Diffusion Models Using Hierarchical Latent PriorsXiao Zhang, Ruoxi Jiang, Rebecca Willett, Michael MaireCVPR 2025
- Not All Parameters Matter: Masking Diffusion Models for Enhancing Generation AbilityLei Wang, Senmao Li, Fei Yang, Jianye Wang et al.CVPR 2025
- AutoDiffusion: Training-Free Optimization of Time Steps and Architectures for Automated Diffusion Model AccelerationLijiang Li, Huixia Li, Xiawu Zheng, Jie Wu et al.ICCV 2023 · 83 citations
