SC2023Top-tier venue
Hanayo: Harnessing Wave-like Pipeline Parallelism for Enhanced Large Model Training Efficiency
Ziming Liu, Shenggan Cheng, Haotian Zhou, Yang You
Abstract
Large-scale language models have become increasingly challenging and expensive to train. Among various methods addressing this issue, Pipeline Parallelism has been widely employed to accommodate massive model weights within limited GPU memory. This paper introduces Hanayo, a wave-like pipeline parallelism strategy that boasts a concise structure and practical applicability, alongside a high-performance pipeline execution runtime to tackle the challenges of pipeline strategy implementation. Hanayo mitigates the issues of pipeline bubbles and excessive memory consumption prevalent in existing schemes, without resorting to model duplicates as in Chimera. Our evaluation, conducted on four distinct computing clusters and involving both GPT-like and BERT-like architectures with up to 32 GPUs, demonstrates up to a 30.4 % increase in throughput compared to the state-of-the-art approach.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c893ec25-841a-4b5f-bfff-af99c672aafdCited by top-tier papers21
- AdaPipe: Optimizing Pipeline Parallelism with Adaptive Recomputation and PartitioningZhenbo Sun, Huanqi Cao, Yuanwei Wang, Guanyu Feng et al.ASPLOS 2024 · 28 citations
- Pipeline Parallelism with Controllable MemoryPenghui Qi, Xinyi Wan, Nyamdavaa Amar, Min LinNeurIPS 2024 · 25 citations
- CrossPipe: Towards Optimal Pipeline Schedules for Cross-Datacenter TrainingTiancheng Chen, Ales Kubicek, Langwen Huang, Torsten HoeflerUSENIX ATC 2025 · 20 citations
- ReCycle: Resilient Training of Large DNNs using Pipeline AdaptationSwapnil Gandhi, Mark Zhao, Athinagoras Skiadopoulos, Christos KozyrakisSOSP 2024 · 13 citations
- GraphPipe: Improving Performance and Scalability of DNN Training with Graph Pipeline ParallelismByungsoo Jeon, Mengdi Wu, Shiyi Cao, Sunghyun Kim et al.ASPLOS 2025 · 10 citations
Builds on11
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu et al.ICLR 2022 · 4,966 citations
- ZeRO: memory optimizations toward training trillion parameter modelsSamyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, Yuxiong HeSC 2020 · 852 citations
Related papers
- Chimera: efficiently training large-scale neural networks with bidirectional pipelinesShigang Li, Torsten HoeflerSC 2021 · 124 citations
- MixPipe: Efficient Bidirectional Pipeline Parallelism for Training Large-Scale ModelsWeigang Zhang, Biyu Zhou, Xuehai Tang, Zhaoxing Wang et al.DAC 2023 · 9 citations
- WeiPipe: Weight Pipeline Parallelism for Communication-Effective Long-Context Large Model TrainingJunfeng Lin, Ziming Liu, Yang You, Jun Wang et al.PPoPP 2025 · 5 citations
- Group-based Interleaved Pipeline Parallelism for Large-scale DNN TrainingPengcheng Yang, Xiaoming Zhang, Wenpeng Zhang, Ming Yang et al.ICLR 2022 · 13 citations
- Efficient large-scale language model training on GPU clusters using megatron-LMDeepak Narayanan, Mohammad Shoeybi, Jared Casper, Patrick LeGresley et al.SC 2021 · 576 citations
