Asynchronous Optimization Methods for Efficient Training of Deep Neural Networks with Guarantees
Vyacheslav Kungurtsev, Malcolm Egan, Bapi Chatterjee, Dan Alistarh
摘要
Asynchronous distributed algorithms are a popular way to reduce synchronization costs in large-scale optimization, and in particular for neural network training. However, for nonsmooth and nonconvex objectives, few convergence guarantees exist beyond cases where closed-form proximal operator solutions are available. As training most popular deep neural networks corresponds to optimizing nonsmooth and nonconvex objectives, there is a pressing need for such convergence guarantees. In this paper, we analyze for the first time the convergence of stochastic asynchronous optimization for this general class of objectives. In particular, we focus on stochastic subgradient methods allowing for block variable partitioning, where the shared model is asynchronously updated by concurrent processes. To this end, we use a probabilistic model which captures key features of real asynchronous scheduling between concurrent processes. Under this model, we establish convergence with probability one to an invariant set for stochastic subgradient methods with momentum. From a practical perspective, one issue with the family of algorithms that we consider is that they are not efficiently supported by machine learning frameworks, which mostly focus on distributed data-parallel strategies. To address this, we propose a new implementation strategy for shared-memory based training of deep neural networks for a partitioned but shared model in single-and multi-GPU settings. Based on this implementation, we achieve on average about 1.2x speedup in comparison to state-of-the-art training methods for popular image classification tasks, without compromising accuracy.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
相关 Paper
- Ordered Momentum for Asynchronous SGDChang-Wei Shi, Yi-Rui Yang, Wu-Jun LiNeurIPS 2024 · 被引用 7 次
- At Stability's Edge: How to Adjust Hyperparameters to Preserve Minima Selection in Asynchronous Training of Neural Networks?Niv Giladi, Mor Shpigel Nacson, Elad Hoffer, Daniel SoudryICLR 2020 · 被引用 25 次
- Ordered Local Momentum for Asynchronous Distributed Learning Under Arbitrary DelaysChang-Wei Shi, Shi-Shang Wang, Wu-Jun LiAAAI 2026
- Elastic Consistency: A Practical Consistency Model for Distributed Stochastic Gradient DescentGiorgi Nadiradze, Ilia Markov, Bapi Chatterjee, Vyacheslav Kungurtsev 等AAAI 2021 · 被引用 10 次
- Gap-Aware Mitigation of Gradient StalenessSaar Barkai, Ido Hakimi, Assaf SchusterICLR 2020 · 被引用 27 次
