A Quadratic Synchronization Rule for Distributed Deep Learning
Xinran Gu, Kaifeng Lyu, Sanjeev Arora, Jingzhao Zhang, Longbo Huang
摘要
In distributed deep learning with data parallelism, synchronizing gradients at each training step can cause a huge communication overhead, especially when many nodes work together to train large models. Local gradient methods, such as Local SGD, address this issue by allowing workers to compute locally for steps without synchronizing with others, hence reducing communication frequency. While has been viewed as a hyperparameter to trade optimization efficiency for communication cost, recent research indicates that setting a proper value can lead to generalization improvement. Yet, selecting a proper is elusive. This work proposes a theory-grounded method for determining , named the Quadratic Synchronization Rule (QSR), which recommends dynamically setting in proportion to as the learning rate decays over time. Extensive ImageNet experiments on ResNet and ViT show that local gradient methods with QSR consistently improve the test accuracy over other synchronization strategies. Compared with the standard data parallel training, QSR enables Local AdamW on ViT-B to cut the training time on 16 or 64 GPUs down from 26.7 to 20.2 hours or from 8.6 to 5.5 hours and, at the same time, achieves or higher top-1 validation accuracy.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Understanding Outer Optimizers in Local SGD: Learning Rates, Momentum, and AccelerationAhmed Khaled, Satyen Kale, Arthur Douillard, Chi Jin 等NeurIPS 2025 · 被引用 7 次
- Adam Reduces a Unique Form of Sharpness: Theoretical Insights Near the Minimizer ManifoldXinghan Li, Haodong Wen, Kaifeng LyuNeurIPS 2025 · 被引用 6 次
- On The Surprising Effectiveness of a Single Global Merging in Decentralized LearningTongtian Zhu, Tianyu Zhang, Mingze Wang, Zhanpeng Zhou 等ICLR 2026 · 被引用 2 次
它引用的顶会 Paper48
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer 等CVPR 2022 · 被引用 6,782 次
- SCAFFOLD: Stochastic Controlled Averaging for Federated LearningSai Praneeth Karimireddy, Satyen Kale, Mehryar Mohri, Sashank J. Reddi 等ICML 2020 · 被引用 3,875 次
相关 Paper
- Don't Use Large Mini-batches, Use Local SGDTao Lin, Sebastian U. Stich, Kumar Kshitij Patel, Martin JaggiICLR 2020 · 被引用 462 次
- A2CiD2: Accelerating Asynchronous Communication in Decentralized Deep LearningAdel Nabli, Eugene Belilovsky, Edouard OyallonNeurIPS 2023 · 被引用 12 次
- DES-LOC: Desynced Low Communication Adaptive Optimizers for Foundation ModelsAlex Iacob, Lorenzo Sani, Mher Safaryan, Paris Giampouras 等ICLR 2026 · 被引用 2 次
- Ordered Local Momentum for Asynchronous Distributed Learning Under Arbitrary DelaysChang-Wei Shi, Shi-Shang Wang, Wu-Jun LiAAAI 2026
- SlowMo: Improving Communication-Efficient Distributed SGD with Slow MomentumJianyu Wang, Vinayak Tantia, Nicolas Ballas, Michael G. RabbatICLR 2020 · 被引用 220 次
