At Stability's Edge: How to Adjust Hyperparameters to Preserve Minima Selection in Asynchronous Training of Neural Networks?
Niv Giladi, Mor Shpigel Nacson, Elad Hoffer, Daniel Soudry
摘要
Background: Recent developments have made it possible to accelerate neural networks training significantly using large batch sizes and data parallelism. Training in an asynchronous fashion, where delay occurs, can make training even more scalable. However, asynchronous training has its pitfalls, mainly a degradation in generalization, even after convergence of the algorithm. This gap remains not well understood, as theoretical analysis so far mainly focused on the convergence rate of asynchronous methods. Contributions: We examine asynchronous training from the perspective of dynamical stability. We find that the degree of delay interacts with the learning rate, to change the set of minima accessible by an asynchronous stochastic gradient descent algorithm. We derive closed-form rules on how the learning rate could be changed, while keeping the accessible set the same. Specifically, for high delay values, we find that the learning rate should be kept inversely proportional to the delay. We then extend this analysis to include momentum. We find momentum should be either turned off, or modified to improve training stability. We provide empirical experiments to validate our theoretical findings.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Stochastic Training is Not Necessary for GeneralizationJonas Geiping, Micah Goldblum, Phillip Pope, Michael Moeller 等ICLR 2022 · 被引用 83 次
- On Linear Stability of SGD and Input-Smoothness of Neural NetworksChao Ma, Lexing YingNeurIPS 2021 · 被引用 73 次
- DoCoFL: Downlink Compression for Cross-Device Federated LearningRon Dorfman, Shay Vargaftik, Yaniv Ben-Itzhak, Kfir Yehuda LevyICML 2023 · 被引用 38 次
- Second-order regression models exhibit progressive sharpening to the edge of stabilityAtish Agarwala, Fabian Pedregosa, Jeffrey PenningtonICML 2023 · 被引用 37 次
- Gradient Descent on Neural Networks Typically Occurs at the Edge of StabilityJeremy Cohen, Simran Kaur, Yuanzhi Li, J. Zico Kolter 等ICLR 2021 · 被引用 22 次
相关 Paper
- Asynchronous Optimization Methods for Efficient Training of Deep Neural Networks with GuaranteesVyacheslav Kungurtsev, Malcolm Egan, Bapi Chatterjee, Dan AlistarhAAAI 2021 · 被引用 4 次
- Gap-Aware Mitigation of Gradient StalenessSaar Barkai, Ido Hakimi, Assaf SchusterICLR 2020 · 被引用 27 次
- Stability and Generalization of Asynchronous SGD: Sharper Bounds Beyond Lipschitz and SmoothnessXiaoge Deng, Tao Sun, Shengwei Li, Dongsheng Li 等NeurIPS 2024 · 被引用 3 次
- Ordered Momentum for Asynchronous SGDChang-Wei Shi, Yi-Rui Yang, Wu-Jun LiNeurIPS 2024 · 被引用 7 次
- The Marginal Value of Momentum for Small Learning Rate SGDRunzhe Wang, Sadhika Malladi, Tianhao Wang, Kaifeng Lyu 等ICLR 2024 · 被引用 14 次
