BPipe: Memory-Balanced Pipeline Parallelism for Training Large Language Models
Taebum Kim, Hyoungjoo Kim, Gyeong-In Yu, Byung-Gon Chun
Abstract
Pipeline parallelism is a key technique for training large language models within GPU clusters. However, it often leads to a memory imbalance problem, where certain GPUs face high memory pressure while others underutilize their capacity. This imbalance results in suboptimal training performance, even when the overall GPU memory capacity is sufficient for more efficient setups. To address this inefficiency, we propose BPIPE, a novel approach for achieving memory balance in pipeline parallelism. BPIPE employs an activation balancing method to transfer intermediate activations between GPUs during training, enabling all GPUs to utilize comparable amounts of memory. With balanced memory utilization, BPIPE enhances the training efficiency of large language models like GPT-3 by eliminating redundant recomputations or increasing the microbatch size. Our evaluation conducted on 48 A100 GPUs across six nodes interconnected with HDR InfiniBand shows that BPIPE accelerates the training of GPT-3 96B and GPT-3 134B models by 1.25x-2.17x compared to Megatron-LM, a stateof-the-art framework for training large language models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 15f3df5f-9ec8-4fb0-a1b3-5952f3a76a79Cited by top-tier papers19
- AdaPipe: Optimizing Pipeline Parallelism with Adaptive Recomputation and PartitioningZhenbo Sun, Huanqi Cao, Yuanwei Wang, Guanyu Feng et al.ASPLOS 2024 · 28 citations
- Pipeline Parallelism with Controllable MemoryPenghui Qi, Xinyi Wan, Nyamdavaa Amar, Min LinNeurIPS 2024 · 25 citations
- CrossPipe: Towards Optimal Pipeline Schedules for Cross-Datacenter TrainingTiancheng Chen, Ales Kubicek, Langwen Huang, Torsten HoeflerUSENIX ATC 2025 · 20 citations
- DropBP: Accelerating Fine-Tuning of Large Language Models by Dropping Backward PropagationSunghyeon Woo, Baeseong Park, Byeongwook Kim, Minjung Jo et al.NeurIPS 2024 · 13 citations
- ReCycle: Resilient Training of Large DNNs using Pipeline AdaptationSwapnil Gandhi, Mark Zhao, Athinagoras Skiadopoulos, Christos KozyrakisSOSP 2024 · 13 citations
Builds on12
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Efficient large-scale language model training on GPU clusters using megatron-LMDeepak Narayanan, Mohammad Shoeybi, Jared Casper, Patrick LeGresley et al.SC 2021 · 576 citations
- Memory-Efficient Pipeline-Parallel DNN TrainingDeepak Narayanan, Amar Phanishayee, Kaiyu Shi, Xie Chen et al.ICML 2021 · 283 citations
- ZeRO-infinity: breaking the GPU memory wall for extreme scale deep learningSamyam Rajbhandari, Olatunji Ruwase, Jeff Rasley, Shaden Smith et al.SC 2021 · 254 citations
Related papers
- WeiPipe: Weight Pipeline Parallelism for Communication-Effective Long-Context Large Model TrainingJunfeng Lin, Ziming Liu, Yang You, Jun Wang et al.PPoPP 2025 · 5 citations
- SlimPipe: Memory-Thrifty and Efficient Pipeline Parallelism for Long-Context LLM TrainingZhouyang Li, Yuliang Liu, Wei Zhang, Tailing Yuan et al.SC 2025 · 5 citations
- Mario: Near Zero-cost Activation Checkpointing in Pipeline ParallelismWeijian Liu, Mingzhen Li, Guangming Tan, Weile JiaPPoPP 2025 · 4 citations
- TawPipe: Topology-Aware Weight Pipeline Parallelism for Accelerating Long-Context Large Models TrainingHouming Wu, Ling ChenAAAI 2026
- PipeOffload: Improving Scalability of Pipeline Parallelism with Memory OptimizationXinyi Wan, Penghui Qi, Guangxing Huang, Min Lin et al.ICML 2025
