Nesterov Method for Asynchronous Pipeline Parallel Optimization
Thalaiyasingam Ajanthan, Sameera Ramasinghe, Yan Zuo, Gil Avraham, Alexander Long
摘要
Pipeline Parallelism (PP) enables large neural network training on small, interconnected devices by splitting the model into multiple stages. To maximize pipeline utilization, asynchronous optimization is appealing as it offers 100% pipeline utilization by construction. However, it is inherently challenging as the weights and gradients are no longer synchronized, leading to stale (or delayed) gradients. To alleviate this, we introduce a variant of Nesterov Accelerated Gradient (NAG) for asynchronous optimization in PP. Specifically, we modify the look-ahead step in NAG to effectively address the staleness in gradients. We theoretically prove that our approach converges at a sublinear rate in the presence of fixed delay in gradients. Our experiments on large-scale language modelling tasks using decoder-only architectures with up to 1B parameters, demonstrate that our approach significantly outperforms existing asynchronous methods, even surpassing the synchronous baseline. †
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Mitigating Staleness in Asynchronous Pipeline Parallelism via Basis RotationHyunji Jung, Sungbin Shin, Namhoon LeeICML 2026 · 被引用 2 次
- AMDP: Asynchronous Multi-Directional Pipeline Parallelism for Large-Scale Models TrainingLing Chen, Houming Wu, Wenjie YuICML 2026 · 被引用 1 次
- PASO: Step Parallel Stochastic OptimizationJianrong Lu, Zhuoya Gu, Haobo Li, Zhiyu Zhu 等ICML 2026 · 被引用 1 次
- One-Step Gradient Delay is Not a Barrier for Large-Scale Asynchronous Pipeline Parallel LLM PretrainingPhilip Zmushko, Egor Petrov, Nursultan Abdullaev, Khrushchev Mikhail 等ICML 2026
- Unextractable Protocol Models: Collaborative Training and Inference without Weight MaterializationAlexander Long, Chamin Hewa Koneputugodage, Thalaiyasingam Ajanthan, Yan Zuo 等NeurIPS 2025
它引用的顶会 Paper7
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Efficient large-scale language model training on GPU clusters using megatron-LMDeepak Narayanan, Mohammad Shoeybi, Jared Casper, Patrick LeGresley 等SC 2021 · 被引用 576 次
- Memory-Efficient Pipeline-Parallel DNN TrainingDeepak Narayanan, Amar Phanishayee, Kaiyu Shi, Xie Chen 等ICML 2021 · 被引用 283 次
- Asynchronous SGD Beats Minibatch SGD Under Arbitrary DelaysKonstantin Mishchenko, Francis R. Bach, Mathieu Even, Blake E. WoodworthNeurIPS 2022 · 被引用 95 次
- Towards Crowdsourced Training of Large Neural Networks using Decentralized Mixture-of-ExpertsMax Ryabinin, Anton GusevNeurIPS 2020 · 被引用 71 次
相关 Paper
- Efficient Pipeline Planning for Expedited Distributed DNN TrainingZiyue Luo, Xiaodong Yi, Guoping Long, Shiqing Fan 等INFOCOM 2022 · 被引用 19 次
- WeiPipe: Weight Pipeline Parallelism for Communication-Effective Long-Context Large Model TrainingJunfeng Lin, Ziming Liu, Yang You, Jun Wang 等PPoPP 2025 · 被引用 5 次
- At Stability's Edge: How to Adjust Hyperparameters to Preserve Minima Selection in Asynchronous Training of Neural Networks?Niv Giladi, Mor Shpigel Nacson, Elad Hoffer, Daniel SoudryICLR 2020 · 被引用 25 次
- PipeGCN: Efficient Full-Graph Training of Graph Convolutional Networks with Pipelined Feature CommunicationCheng Wan, Youjie Li, Cameron R. Wolfe, Anastasios Kyrillidis 等ICLR 2022 · 被引用 89 次
- Accumulated Decoupled Learning with Gradient Staleness Mitigation for Convolutional Neural NetworksHuiping Zhuang, Zhenyu Weng, Fulin Luo, Kar-Ann Toj 等ICML 2021 · 被引用 6 次
