Heterogeneous Decentralized Diffusion Models
Zhiying Jiang, Raihan Seraj, Marcos Villagra, Bidhan Roy
Abstract
Training state-of-the-art diffusion models requires massive computational resources concentrated in tightly-coupled clusters, fundamentally limiting participation to well-resourced institutions. While Decentralized Diffusion Models (DDM) enable training multiple experts in isolation, existing approaches require 1176 GPU-days and homogeneous training objectives across all experts. We present an efficient framework that dramatically reduces resource requirements while supporting heterogeneous training objectives. Our approach combines three key contributions: (1) PixArt-'s efficient AdaLN-Single architecture, reducing parameters while maintaining quality; (2) pretrained checkpoint conversion from ImageNet-DDPM to Flow Matching objectives, accelerating convergence and enabling initialization without objective-specific pretraining; and (3) a training-free inference conversion framework that unifies heterogeneous expert predictions (DDPM and Flow Matching) into a common velocity space without any retraining. Experiments on LAION-Aesthetics demonstrate that our decentralized approach achieves comparative results with 16 compute reduction (72 vs 1176 GPU-days) and 14 data reduction (11M vs 158M images). Our heterogeneous variant mixing DDPM and Flow Matching experts exhibits complementary specialization patterns, improving generation diversity and texture quality despite modest FID increases. By eliminating synchronization requirements and enabling arbitrary objective combinations, our framework democratizes large-scale generative model training, allowing contributors with diverse resources to participate using consumer GPUs requiring only 20-48GB VRAM.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 070ca375-670a-49cb-a3c1-789bc96915c7Builds on23
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
Related papers
- Decentralized Diffusion ModelsDavid McAllister, Matthew Tancik, Jiaming Song, Angjoo KanazawaCVPR 2025
- PixArt-α: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image SynthesisJunsong Chen, Jincheng Yu, Chongjian Ge, Lewei Yao et al.ICLR 2024 · 831 citations
- Learning Few-Step Diffusion Models by Trajectory Distribution MatchingYihong Luo, Tianyang Hu, Jiacheng Sun, Yujun Cai et al.ICCV 2025 · 3 citations
- FreePIH: Training-Free Painterly Image Harmonization with Diffusion ModelRuibin Li, Jingcai Guo, Qihua Zhou, Song GuoACM MM 2024 · 2 citations
- Patch Diffusion: Faster and More Data-Efficient Training of Diffusion ModelsZhendong Wang, Yifan Jiang, Huangjie Zheng, Peihao Wang et al.NeurIPS 2023 · 205 citations
