Practical Performance Guarantees for Pipelined DNN Inference
Aaron Archer, Matthew Fahrbach, Kuikui Liu, Prakash Prabhu
Abstract
We optimize pipeline parallelism for deep neural network (DNN) inference by partitioning model graphs into stages and minimizing the running time of the bottleneck stage, including communication. We give practical and effective algorithms for this NP-hard problem, but our emphasis is on tackling the practitioner's dilemma of deciding when a solution is good enough. To this end, we design novel mixed-integer programming (MIP) relaxations for proving lower bounds. Applying these methods to a diverse testbed of 369 production models, for , we empirically show that these lower bounds are strong enough to be useful in practice. Our lower bounds are substantially stronger than standard combinatorial bounds. For example, evaluated via geometric means across a production testbed with pipeline stages, our MIP formulations raise the lower bound from 0.4598 to 0.9452, expressed as a fraction of the best partition found. In other words, our improved lower bounds close the optimality gap by a factor of 9.855x.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d2a0f00f-0e2a-4189-94fe-1c8cd7656b1eCited by top-tier papers1
Ask how each one uses itBuilds on8
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- Memory-Efficient Pipeline-Parallel DNN TrainingDeepak Narayanan, Amar Phanishayee, Kaiyu Shi, Xie Chen et al.ICML 2021 · 283 citations
- TeraPipe: Token-Level Pipeline Parallelism for Training Large-Scale Language ModelsZhuohan Li, Siyuan Zhuang, Shiyuan Guo, Danyang Zhuo et al.ICML 2021 · 160 citations
- Decentralized Training of Foundation Models in Heterogeneous EnvironmentsBinhang Yuan, Yongjun He, Jared Davis, Tianyi Zhang et al.NeurIPS 2022 · 157 citations
- Reinforced Genetic Algorithm Learning for Optimizing Computation GraphsAditya Paliwal, Felix Gimeno, Vinod Nair, Yujia Li et al.ICLR 2020 · 70 citations
Related papers
- Aceso: Efficient Parallel DNN Training through Iterative Bottleneck AlleviationGuodong Liu, Youshan Miao, Zhiqi Lin, Xiaoxiang Shi et al.EuroSys 2024 · 16 citations
- Partitioned Scheduling and Parallelism Assignment for Real-Time DNN Inference Tasks on Multi-TPUBinqi Sun, Tomasz Kloda, Chu-Ge Wu, Marco CaccamoDAC 2024 · 8 citations
- GraphPipe: Improving Performance and Scalability of DNN Training with Graph Pipeline ParallelismByungsoo Jeon, Mengdi Wu, Shiyi Cao, Sunghyun Kim et al.ASPLOS 2025 · 10 citations
- Efficient Algorithms for Device Placement of DNN Graph OperatorsJakub Tarnawski, Amar Phanishayee, Nikhil R. Devanur, Divya Mahajan et al.NeurIPS 2020 · 84 citations
- PPipe: Efficient Video Analytics Serving on Heterogeneous GPU Clusters via Pool-Based Pipeline ParallelismZ. Jonny Kong, Qiang Xu, Y. Charlie HuUSENIX ATC 2025 · 4 citations
