Efficient Algorithms for Device Placement of DNN Graph Operators
Jakub Tarnawski, Amar Phanishayee, Nikhil R. Devanur, Divya Mahajan, Fanny Nina Paravecino
摘要
Modern machine learning workloads use large models, with complex structures, that are very expensive to execute. The devices that execute complex models are becoming increasingly heterogeneous as we see a flourishing of domain-specific accelerators being offered as hardware accelerators in addition to CPUs. These trends necessitate distributing the workload across multiple devices. Recent work has shown that significant gains can be obtained with model parallelism, i.e, partitioning a neural network's computational graph onto multiple devices. In particular, this form of parallelism assumes a pipeline of devices, which is fed a stream of samples and yields high throughput for training and inference of DNNs. However, for such settings (large models and multiple heterogeneous devices), we require automated algorithms and toolchains that can partition the ML workload across devices. In this paper, we identify and isolate the structured optimization problem at the core of device placement of DNN operators, for both inference and training, especially in modern pipelined settings. We then provide algorithms that solve this problem to optimality. We demonstrate the applicability and efficiency of our approaches using several contemporary DNN computation graphs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper16
- Efficient large-scale language model training on GPU clusters using megatron-LMDeepak Narayanan, Mohammad Shoeybi, Jared Casper, Patrick LeGresley 等SC 2021 · 被引用 576 次
- TopoOpt: Co-optimizing Network Topology and Parallelization Strategy for Distributed Training JobsWeiyang Wang, Moein Khazraee, Zhizhen Zhong, Manya Ghobadi 等NSDI 2023 · 被引用 215 次
- Decentralized Training of Foundation Models in Heterogeneous EnvironmentsBinhang Yuan, Yongjun He, Jared Davis, Tianyi Zhang 等NeurIPS 2022 · 被引用 157 次
- NeuPIMs: NPU-PIM Heterogeneous Acceleration for Batched LLM InferencingGuseul Heo, Sangyeop Lee, Jaehong Cho, Hyunmin Choi 等ASPLOS 2024 · 被引用 121 次
- DFX: A Low-latency Multi-FPGA Appliance for Accelerating Transformer-based Text GenerationSeongmin Hong, Seungjae Moon, Junsoo Kim, Sungjae Lee 等MICRO 2022 · 被引用 107 次
相关 Paper
- Efficient Pipeline Planning for Expedited Distributed DNN TrainingZiyue Luo, Xiaodong Yi, Guoping Long, Shiqing Fan 等INFOCOM 2022 · 被引用 19 次
- GraphPipe: Improving Performance and Scalability of DNN Training with Graph Pipeline ParallelismByungsoo Jeon, Mengdi Wu, Shiyi Cao, Sunghyun Kim 等ASPLOS 2025 · 被引用 10 次
- Piper: Multidimensional Planner for DNN ParallelizationJakub Tarnawski, Deepak Narayanan, Amar PhanishayeeNeurIPS 2021 · 被引用 82 次
- Tessel: Boosting Distributed Execution of Large DNN Models via Flexible Schedule SearchZhiqi Lin, Youshan Miao, Guanbin Xu, Cheng Li 等HPCA 2024 · 被引用 6 次
- QuiltNet: efficient deep learning inference on multi-chip accelerators using model partitioningJongho Park, Hyukjun Kwon, Seowoo Kim, Junyoung Lee 等DAC 2022 · 被引用 6 次
