Pfeife: Automatic Pipeline Parallelism for PyTorch
Ho Young Jhoo, Chung-Kil Hur, Nuno P. Lopes
Abstract
The memory requirements of machine learning (ML) models has been growing quickly. However, the memory capacity of GPUs has not kept pace. Despite significant research on reducing the memory usage of ML models, the larger models do not fit in a single device. A popular solution to the memory capacity issue is to use multiple devices in parallel. In this paper, we focus on a particular form of parallelism called pipelining, as it offers a good balance between cost and performance for many ML models. We present Pfeife, the first tool that integrates with PyTorch to provide automatic pipelining of ML models. Pfeife intercepts the execution of models and parallelizes them transparently, requiring no manual work. We show that Pfeife can execute large models that would otherwise not run due to not fitting in a single device. Moreover, Pfeife can pipeline non-sequential models such as Stable Diffusion, which are not supported by existing pipelining parallelism tools. Pfeife outperforms state-of-the-art tools by up to 22%.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext aa7b48a1-4f53-4ec8-a00a-4737a212e151Builds on12
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann et al.ICLR 2024 · 4,569 citations
- ZeRO: memory optimizations toward training trillion parameter modelsSamyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, Yuxiong HeSC 2020 · 852 citations
- Scaling Vision TransformersXiaohua Zhai, Alexander Kolesnikov, Neil Houlsby, Lucas BeyerCVPR 2022 · 767 citations
- PyTorch 2: Faster Machine Learning Through Dynamic Python Bytecode Transformation and Graph CompilationJason Ansel, Edward Z. Yang, Horace He, Natalia Gimelshein et al.ASPLOS 2024 · 693 citations
Related papers
- Elastic Averaging for Efficient Pipelined DNN TrainingZihao Chen, Chen Xu, Weining Qian, Aoying ZhouPPoPP 2023 · 10 citations
- PipeFusion: Patch-level Pipeline Parallelism for Diffusion Transformers InferenceJiarui Fang, Jinzhe Pan, Aoyu Li, Xibo Sun et al.NeurIPS 2025 · 36 citations
- Fine-tuning giant neural networks on commodity hardware with automatic pipeline model parallelismSaar Eliad, Ido Hakimi, Alon De Jagger, Mark Silberstein et al.USENIX ATC 2021 · 24 citations
- AsyncDiff: Parallelizing Diffusion Models by Asynchronous DenoisingZigeng Chen, Xinyin Ma, Gongfan Fang, Zhenxiong Tan et al.NeurIPS 2024 · 33 citations
- Efficient Algorithms for Device Placement of DNN Graph OperatorsJakub Tarnawski, Amar Phanishayee, Nikhil R. Devanur, Divya Mahajan et al.NeurIPS 2020 · 84 citations
