DropCompute: simple and more robust distributed synchronous training via compute variance reduction
Niv Giladi, Shahar Gottlieb, Moran Shkolnik, Asaf Karnieli, Ron Banner, Elad Hoffer, Kfir Y. Levy, Daniel Soudry
摘要
Background. Distributed training is essential for large scale training of deep neural networks (DNNs). The dominant methods for large scale DNN training are synchronous (e.g. All-Reduce), but these require waiting for all workers in each step. Thus, these methods are limited by the delays caused by straggling workers. Results. We study a typical scenario in which workers are straggling due to variability in compute time. We find an analytical relation between compute time properties and scalability limitations, caused by such straggling workers. With these findings, we propose a simple yet effective decentralized method to reduce the variation among workers and thus improve the robustness of synchronous training. This method can be integrated with the widely used All-Reduce. Our findings are validated on large-scale training tasks using 200 Gaudi Accelerators. A reference implementation 2 is provided. * Equal contribution 2 https://github.com/paper-submissions/dropcompute 37th Conference on Neural Information Processing Systems (NeurIPS 2023).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Understanding Stragglers in Large Model Training Using What-if AnalysisJinkun Lin, Ziheng Jiang, Zuquan Song, Sida Zhao 等OSDI 2025 · 被引用 23 次
- From Promise to Practice: Realizing High-performance Decentralized TrainingZesen Wang, Jiaojiao Zhang, Xuyang Wu, Mikael JohanssonICLR 2025
- Trajectory-Aware Certified Decentralized Unlearning via SGD StabilityHengliang Wu, Jiale Yang, Shuzhen Chen, Di Wang 等ICML 2026
它引用的顶会 Paper10
- EfficientNetV2: Smaller Models and Faster TrainingMingxing Tan, Quoc V. LeICML 2021 · 被引用 4,239 次
- Large Batch Optimization for Deep Learning: Training BERT in 76 minutesYang You, Jing Li, Sashank J. Reddi, Jonathan Hseu 等ICLR 2020 · 被引用 1,170 次
- ZeRO: memory optimizations toward training trillion parameter modelsSamyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, Yuxiong HeSC 2020 · 被引用 852 次
- Efficient large-scale language model training on GPU clusters using megatron-LMDeepak Narayanan, Mohammad Shoeybi, Jared Casper, Patrick LeGresley 等SC 2021 · 被引用 576 次
- Don't Use Large Mini-batches, Use Local SGDTao Lin, Sebastian U. Stich, Kumar Kshitij Patel, Martin JaggiICLR 2020 · 被引用 462 次
相关 Paper
- Near-Optimal Topology-adaptive Parameter Synchronization in Distributed DNN TrainingZhe Zhang, Chuan Wu, Zongpeng LiINFOCOM 2021 · 被引用 14 次
- Gsyn: Reducing Staleness and Communication Waiting via Grouping-based Synchronization for Distributed Deep LearningYijun Li, Jiawei Huang, Zhaoyi Li, Jingling Liu 等INFOCOM 2024 · 被引用 2 次
- Gap-Aware Mitigation of Gradient StalenessSaar Barkai, Ido Hakimi, Assaf SchusterICLR 2020 · 被引用 27 次
- OptiReduce: Resilient and Tail-Optimal AllReduce for Distributed Deep Learning in the CloudErtza Warraich, Omer Shabtai, Khalid Manaa, Shay Vargaftik 等NSDI 2025
- Decentralized Diffusion ModelsDavid McAllister, Matthew Tancik, Jiaming Song, Angjoo KanazawaCVPR 2025
