Nesterov Method for Asynchronous Pipeline Parallel Optimization
Thalaiyasingam Ajanthan, Sameera Ramasinghe, Yan Zuo, Gil Avraham, Alexander Long
Abstract
Pipeline Parallelism (PP) enables large neural network training on small, interconnected devices by splitting the model into multiple stages. To maximize pipeline utilization, asynchronous optimization is appealing as it offers 100% pipeline utilization by construction. However, it is inherently challenging as the weights and gradients are no longer synchronized, leading to stale (or delayed) gradients. To alleviate this, we introduce a variant of Nesterov Accelerated Gradient (NAG) for asynchronous optimization in PP. Specifically, we modify the look-ahead step in NAG to effectively address the staleness in gradients. We theoretically prove that our approach converges at a sublinear rate in the presence of fixed delay in gradients. Our experiments on large-scale language modelling tasks using decoder-only architectures with up to 1B parameters, demonstrate that our approach significantly outperforms existing asynchronous methods, even surpassing the synchronous baseline. †
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- Mitigating Staleness in Asynchronous Pipeline Parallelism via Basis RotationHyunji Jung, Sungbin Shin, Namhoon LeeICML 2026 · 2 citations
- AMDP: Asynchronous Multi-Directional Pipeline Parallelism for Large-Scale Models TrainingLing Chen, Houming Wu, Wenjie YuICML 2026 · 1 citation
- PASO: Step Parallel Stochastic OptimizationJianrong Lu, Zhuoya Gu, Haobo Li, Zhiyu Zhu et al.ICML 2026 · 1 citation
- One-Step Gradient Delay is Not a Barrier for Large-Scale Asynchronous Pipeline Parallel LLM PretrainingPhilip Zmushko, Egor Petrov, Nursultan Abdullaev, Khrushchev Mikhail et al.ICML 2026
- Unextractable Protocol Models: Collaborative Training and Inference without Weight MaterializationAlexander Long, Chamin Hewa Koneputugodage, Thalaiyasingam Ajanthan, Yan Zuo et al.NeurIPS 2025
Builds on7
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Efficient large-scale language model training on GPU clusters using megatron-LMDeepak Narayanan, Mohammad Shoeybi, Jared Casper, Patrick LeGresley et al.SC 2021 · 576 citations
- Memory-Efficient Pipeline-Parallel DNN TrainingDeepak Narayanan, Amar Phanishayee, Kaiyu Shi, Xie Chen et al.ICML 2021 · 283 citations
- Asynchronous SGD Beats Minibatch SGD Under Arbitrary DelaysKonstantin Mishchenko, Francis R. Bach, Mathieu Even, Blake E. WoodworthNeurIPS 2022 · 95 citations
- Towards Crowdsourced Training of Large Neural Networks using Decentralized Mixture-of-ExpertsMax Ryabinin, Anton GusevNeurIPS 2020 · 71 citations
Related papers
- Efficient Pipeline Planning for Expedited Distributed DNN TrainingZiyue Luo, Xiaodong Yi, Guoping Long, Shiqing Fan et al.INFOCOM 2022 · 19 citations
- WeiPipe: Weight Pipeline Parallelism for Communication-Effective Long-Context Large Model TrainingJunfeng Lin, Ziming Liu, Yang You, Jun Wang et al.PPoPP 2025 · 5 citations
- At Stability's Edge: How to Adjust Hyperparameters to Preserve Minima Selection in Asynchronous Training of Neural Networks?Niv Giladi, Mor Shpigel Nacson, Elad Hoffer, Daniel SoudryICLR 2020 · 25 citations
- PipeGCN: Efficient Full-Graph Training of Graph Convolutional Networks with Pipelined Feature CommunicationCheng Wan, Youjie Li, Cameron R. Wolfe, Anastasios Kyrillidis et al.ICLR 2022 · 89 citations
- Accumulated Decoupled Learning with Gradient Staleness Mitigation for Convolutional Neural NetworksHuiping Zhuang, Zhenyu Weng, Fulin Luo, Kar-Ann Toj et al.ICML 2021 · 6 citations
