Zero Bubble (Almost) Pipeline Parallelism
Penghui Qi, Xinyi Wan, Guangxing Huang, Min Lin
摘要
Pipeline parallelism is one of the key components for large-scale distributed training, yet its efficiency suffers from pipeline bubbles which were deemed inevitable. In this work, we introduce a scheduling strategy that, to our knowledge, is the first to successfully achieve zero pipeline bubbles under synchronous training semantics. The key idea behind this improvement is to split the backward computation into two parts, one that computes gradient for the input and another that computes for the parameters. Based on this idea, we handcraft novel pipeline schedules that significantly outperform the baseline methods. We further develop an algorithm that automatically finds an optimal schedule based on specific model configuration and memory limit. Additionally, to truly achieve zero bubble, we introduce a novel technique to bypass synchronizations during the optimizer step. Experimental evaluations show that our method outperforms the 1F1B schedule up to 15% in throughput under a similar memory limit. This number can be further pushed to 30% when the memory constraint is relaxed. We believe our results mark a major step forward in harnessing the true potential of pipeline parallelism. The source code based on Megatron-LM is publicly avaiable at https: //github.com/sail-sg/zero-bubble-pipeline-parallelism .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper27
- mHC: Manifold-Constrained Hyper-ConnectionsZhenda Xie, Yixuan Wei, Huanqi Cao, Chenggang Zhao 等ICML 2026 · 被引用 72 次
- EXP-Bench: Can AI Conduct AI Research Experiments?Patrick Tser Jern Kon, Qiuyi Ding, Jiachen Liu, Xinyi Zhu 等ICLR 2026 · 被引用 35 次
- Rethinking the Trust Region in LLM Reinforcement LearningPenghui Qi, Xiangxin Zhou, Zichen Liu, Tianyu Pang 等ICML 2026 · 被引用 22 次
- MEPipe: Democratizing LLM Training with Memory-Efficient Slice-Level Pipeline Scheduling on Cost-Effective AcceleratorsZhenbo Sun, Shengqi Chen, Yuanwei Wang, Jian Sha 等EuroSys 2025 · 被引用 7 次
- Mycroft: Tracing Dependencies in Collective Communication Towards Reliable LLM TrainingYangtao Deng, Lei Zhang, Qinlong Wang, Xiaoyun Zhi 等SOSP 2025 · 被引用 7 次
它引用的顶会 Paper5
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- ZeRO: memory optimizations toward training trillion parameter modelsSamyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, Yuxiong HeSC 2020 · 被引用 852 次
- Efficient large-scale language model training on GPU clusters using megatron-LMDeepak Narayanan, Mohammad Shoeybi, Jared Casper, Patrick LeGresley 等SC 2021 · 被引用 576 次
- DAPPLE: a pipelined data parallel approach for training large modelsShiqing Fan, Yi Rong, Chen Meng, Zongyan Cao 等PPoPP 2021 · 被引用 224 次
- Alpa: Automating Inter- and Intra-Operator Parallelism for Distributed Deep LearningLianmin Zheng, Zhuohan Li, Hao Zhang, Yonghao Zhuang 等OSDI 2022 · 被引用 75 次
相关 Paper
- Synergistic Tensor and Pipeline ParallelismMengshi Qi, Jiaxuan Peng, Jie M. Zhang, Juan Zhu 等NeurIPS 2025 · 被引用 2 次
- Pipeline Parallelism with Controllable MemoryPenghui Qi, Xinyi Wan, Nyamdavaa Amar, Min LinNeurIPS 2024 · 被引用 25 次
- AdaptPipe: Mitigating Runtime Bubbles via Granularity-Adaptive Scheduling under Memory ConstraintsYumeng Cui, Jessie Hui Wang, Najila Liu, Ling Deng 等INFOCOM 2026
- HelixPipe: Efficient Distributed Training of Long Sequence Transformers with Attention Parallel Pipeline ParallelismGeng Zhang, Shenggan Cheng, Xuanlei Zhao, Ziming Liu 等PPoPP 2026 · 被引用 3 次
- MixPipe: Efficient Bidirectional Pipeline Parallelism for Training Large-Scale ModelsWeigang Zhang, Biyu Zhou, Xuehai Tang, Zhaoxing Wang 等DAC 2023 · 被引用 9 次
