Lune

ICLR2024Top-tier venue

A Quadratic Synchronization Rule for Distributed Deep Learning

Xinran Gu, Kaifeng Lyu, Sanjeev Arora, Jingzhao Zhang, Longbo Huang

2024Year
4Citations
3Top-tier citations

Abstract

In distributed deep learning with data parallelism, synchronizing gradients at each training step can cause a huge communication overhead, especially when many nodes work together to train large models. Local gradient methods, such as Local SGD, address this issue by allowing workers to compute locally for HH steps without synchronizing with others, hence reducing communication frequency. While HH has been viewed as a hyperparameter to trade optimization efficiency for communication cost, recent research indicates that setting a proper HH value can lead to generalization improvement. Yet, selecting a proper HH is elusive. This work proposes a theory-grounded method for determining HH, named the Quadratic Synchronization Rule (QSR), which recommends dynamically setting HH in proportion to 1η2\frac{1}{\eta^2} as the learning rate η\eta decays over time. Extensive ImageNet experiments on ResNet and ViT show that local gradient methods with QSR consistently improve the test accuracy over other synchronization strategies. Compared with the standard data parallel training, QSR enables Local AdamW on ViT-B to cut the training time on 16 or 64 GPUs down from 26.7 to 20.2 hours or from 8.6 to 5.5 hours and, at the same time, achieves 1.16%1.16\% or 0.84%0.84\% higher top-1 validation accuracy.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext da0693b8-0459-4fb9-89dc-e7b793bdf437

Cited by top-tier papers3

Ask how each one uses it

Builds on48

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines