Lune

NeurIPS2023顶会

A2CiD2: Accelerating Asynchronous Communication in Decentralized Deep Learning

Adel Nabli, Eugene Belilovsky, Edouard Oyallon

2023年份
12被引次数

摘要

Distributed training of Deep Learning models has been critical to many recent successes in the field. Current standard methods primarily rely on synchronous centralized algorithms which induce major communication bottlenecks and synchronization locks at scale. Decentralized asynchronous algorithms are emerging as a potential alternative but their practical applicability still lags. In order to mitigate the increase in communication cost that naturally comes with scaling the number of workers, we introduce a principled asynchronous, randomized, gossip-based optimization algorithm which works thanks to a continuous local momentum named A2CiD2\textbf{A}^2\textbf{CiD}^2. Our method allows each worker to continuously process mini-batches without stopping, and run a peer-to-peer averaging routine in parallel, reducing idle time. In addition to inducing a significant communication acceleration at no cost other than adding a local momentum variable, minimal adaptation is required to incorporate A2CiD2\textbf{A}^2\textbf{CiD}^2 to standard asynchronous approaches. Our theoretical analysis proves accelerated rates compared to previous asynchronous decentralized baselines and we empirically show that using our A2CiD2\textbf{A}^2\textbf{CiD}^2 momentum significantly decrease communication costs in poorly connected networks. In particular, we show consistent improvement on the ImageNet dataset using up to 64 asynchronous workers (A100 GPUs) and various communication network topologies.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

lune papers fulltext cb301af9-0373-4fdf-a3ae-e3ce43eb9f34

它引用的顶会 Paper15

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖