Decentralized Diffusion Models
David McAllister, Matthew Tancik, Jiaming Song, Angjoo Kanazawa
摘要
Large-scale AI model training divides work across thousands of GPUs, then synchronizes gradients across them at each step. This incurs a significant network burden that only centralized, monolithic clusters can support, driving up infrastructure costs and straining power systems. We propose Decentralized Diffusion Models, a scalable framework for distributing diffusion model training across independent clusters or datacenters by eliminating the dependence on a centralized, high-bandwidth networking fabric. Our method trains a set of expert diffusion models over partitions of the dataset, each in full isolation from one another. At inference time, the experts ensemble through a lightweight router. We show that the ensemble collectively optimizes the same objective as a single model trained over the whole dataset. This means we can divide the training burden among a number of "compute islands," lowering infrastructure costs and improving resilience to localized GPU failures. Decentralized diffusion models empower researchers to take advantage of smaller, more cost-effective and more readily available compute like on-demand GPU nodes rather than central integrated systems. We conduct extensive experiments on ImageNet and LAION Aesthetics, showing that decentralized diffusion models FLOP-for-FLOP outperform standard diffusion models. We finally scale our approach to 24 billion parameters, demonstrating that high-quality diffusion models can now be trained with just eight individual GPU nodes in less than a week.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Heterogeneous Decentralized Diffusion ModelsZhiying Jiang, Raihan Seraj, Marcos Villagra, Bidhan RoyCVPR 2026 · 被引用 1 次
- Compositional Generative Modeling from Decentralized DataMashrur M. Morshed, Vishnu BoddetiICML 2026
- Mechanisms of Projective Composition of Diffusion ModelsArwen Bradley, Preetum Nakkiran, David Berthelot, James Thornton 等ICML 2025
它引用的顶会 Paper23
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
相关 Paper
- Beyond A Single AI Cluster: A Survey of Decentralized LLM TrainingHaotian Dong, Jingyan Jiang, Rongwei Lu, Jiajun Luo 等EMNLP 2025 · 被引用 2 次
- Patch Diffusion: Faster and More Data-Efficient Training of Diffusion ModelsZhendong Wang, Yifan Jiang, Huangjie Zheng, Peihao Wang 等NeurIPS 2023 · 被引用 205 次
- DropCompute: simple and more robust distributed synchronous training via compute variance reductionNiv Giladi, Shahar Gottlieb, Moran Shkolnik, Asaf Karnieli 等NeurIPS 2023 · 被引用 4 次
- DLion: Decentralized Distributed Deep Learning in Micro-CloudsRankyung Hong, Abhishek ChandraHPDC 2021 · 被引用 20 次
- Consensus Control for Decentralized Deep LearningLingjing Kong, Tao Lin, Anastasia Koloskova, Martin Jaggi 等ICML 2021 · 被引用 100 次
