Elastic Averaging for Efficient Pipelined DNN Training
Zihao Chen, Chen Xu, Weining Qian, Aoying Zhou
Abstract
Nowadays, the size of DNN models has grown rapidly. To train a large model, pipeline parallelism-based frameworks partition the model across GPUs and slice each batch of data into multiple micro-batches. However, pipeline parallelism suffers from a bubble issue and low peak utilization of GPUs. Recent work tries to address the two issues, but fails to exploit the benefit of vanilla pipeline parallelism, i.e., overlapping communication with computation. In this work, we employ an elastic averaging-based framework which explores elastic averaging to add multiple parallel pipelines. To help the framework exploit the advantage of pipeline parallelism while reducing the memory footprints, we propose a schedule, advance forward propagation. Moreover, since the numbers of parallel pipelines and micro-batches are essential to the framework performance, we propose a profiling-based tuning method to automatically determine the settings. We integrate those techniques into a prototype system, namely AvgPipe, based on PyTorch. Our experiments show that Avg-Pipe achieves a 1.7x speedups over state-of-the-art solutions of pipeline parallelism on average.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get accf72c6-8a25-4070-8020-cf58791f8f7cRelated papers
- AdaptPipe: Mitigating Runtime Bubbles via Granularity-Adaptive Scheduling under Memory ConstraintsYumeng Cui, Jessie Hui Wang, Najila Liu, Ling Deng et al.INFOCOM 2026
- EnvPipe: Performance-preserving DNN Training Framework for Saving EnergySangjin Choi, Inhoe Koo, Jeongseob Ahn, Myeongjae Jeon et al.USENIX ATC 2023 · 41 citations
- GraphPipe: Improving Performance and Scalability of DNN Training with Graph Pipeline ParallelismByungsoo Jeon, Mengdi Wu, Shiyi Cao, Sunghyun Kim et al.ASPLOS 2025 · 10 citations
- Efficient Pipeline Planning for Expedited Distributed DNN TrainingZiyue Luo, Xiaodong Yi, Guoping Long, Shiqing Fan et al.INFOCOM 2022 · 19 citations
- Zero Bubble (Almost) Pipeline ParallelismPenghui Qi, Xinyi Wan, Guangxing Huang, Min LinICLR 2024 · 31 citations
