AdaptPipe: Mitigating Runtime Bubbles via Granularity-Adaptive Scheduling under Memory Constraints
Yumeng Cui, Jessie Hui Wang, Najila Liu, Ling Deng, Liang Du, Chuxuan Zeng, Jilong Wang
摘要
Crucially, pipeline parallelism (PP) is an indispensable parallelization strategy for training large models on multiple GPUs. Despite some proposed PP schedules theoretically promising zero bubbles, we highlight that runtime bubbles remain a significant practical hurdle in real-world model training. This discrepancy arises because these idealized schedules presuppose that F (forward), B (backward), and W (weight) operations have uniform and stable execution times. This assumption does not hold in practical scenarios.In this paper, we introduce AdaptPipe, an adaptive pipeline scheduling framework that mitigates runtime bubbles by leveraging our analysis of the characteristics of two types of runtime bubbles. AdaptPipe improves runtime efficiency by opportunistically filling W tasks of appropriate granularity. To achieve this, we develop a prediction method and a notification method to estimate the size of different runtime bubbles. Furthermore, we design a forward-backward (FB) scheme, namely nF –nB, which aims to provide AdaptPipe with sufficient schedulable W tasks to handle continuous or large bubbles while still adhering to memory constraints.Our experiments on dense models and sparse MoE models show that AdaptPipe improves throughput by 0.6%-13.8% over state-of-the-art static pipeline schedules while maintaining comparable activation memory usage.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- Zero Bubble (Almost) Pipeline ParallelismPenghui Qi, Xinyi Wan, Guangxing Huang, Min LinICLR 2024 · 被引用 31 次
- Elastic Averaging for Efficient Pipelined DNN TrainingZihao Chen, Chen Xu, Weining Qian, Aoying ZhouPPoPP 2023 · 被引用 10 次
- Synergistic Tensor and Pipeline ParallelismMengshi Qi, Jiaxuan Peng, Jie M. Zhang, Juan Zhu 等NeurIPS 2025 · 被引用 2 次
- Pipeline Parallelism with Controllable MemoryPenghui Qi, Xinyi Wan, Nyamdavaa Amar, Min LinNeurIPS 2024 · 被引用 25 次
- WeiPipe: Weight Pipeline Parallelism for Communication-Effective Long-Context Large Model TrainingJunfeng Lin, Ziming Liu, Yang You, Jun Wang 等PPoPP 2025 · 被引用 5 次
