Layerweaver: Maximizing Resource Utilization of Neural Processing Units via Layer-Wise Scheduling
Young H. Oh, Seonghak Kim, Yunho Jin, Sam Son, Jonghyun Bae, Jongsung Lee, Yeonhong Park, Dong Uk Kim, Tae Jun Ham, Jae W. Lee
摘要
To meet surging demands for deep learning inference services, many cloud computing vendors employ high-performance specialized accelerators, called neural processing units (NPUs). One important challenge for effective use of NPUs is to achieve high resource utilization over a wide spectrum of deep neural network (DNN) models with diverse arithmetic intensities. There is often an intrinsic mismatch between the compute-to-memory bandwidth ratio of an NPU and the arithmetic intensity of the model it executes, leading to under-utilization of either compute resources or memory bandwidth. Ideally, we want to saturate both compute TOP/s and DRAM bandwidth to achieve high system throughput. Thus, we propose Layerweaver, an inference serving system with a novel multi-model time-multiplexing scheduler for NPUs. Layerweaver reduces the temporal waste of computation resources by interweaving layer execution of multiple different models with opposing characteristics: compute-intensive and memory-intensive. Layerweaver hides the memory time of a memory-intensive model by overlapping it with the relatively long computation time of a compute-intensive model, thereby minimizing the idle time of the computation units waiting for off-chip data transfers. For a two-model serving scenario of batch 1 with 16 different pairs of compute- and memory-intensive models, Layerweaver improves the temporal utilization of computation units and memory channels by 44.0% and 28.7%, respectively, to increase the system throughput by 60.1% on average, over the baseline executing one model at a time.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Behemoth: A Flash-centric Training Accelerator for Extreme-scale DNNsShine Kim, Yunho Jin, Gina Sohn, Jonghyun Bae 等FAST 2021 · 被引用 44 次
- MoCA: Memory-Centric, Adaptive Execution for Multi-Tenant Deep Neural NetworksSeah Kim, Hasan Genc, Vadim Vadimovich Nikiforov, Krste Asanovic 等HPCA 2023 · 被引用 36 次
- V10: Hardware-Assisted NPU Multi-tenancy for Improved Resource Utilization and FairnessYuqi Xue, Yiqi Liu, Lifeng Nai, Jian HuangISCA 2023 · 被引用 21 次
- Sparse-DySta: Sparsity-Aware Dynamic and Static Scheduling for Sparse Multi-DNN WorkloadsHongxiang Fan, Stylianos I. Venieris, Alexandros Kouris, Nicholas D. LaneMICRO 2023 · 被引用 12 次
- Hardware-Assisted Virtualization of Neural Processing Units for Cloud PlatformsYuqi Xue, Yiqi Liu, Lifeng Nai, Jian HuangMICRO 2024 · 被引用 10 次
它引用的顶会 Paper3
- A3: Accelerating Attention Mechanisms in Neural Networks with ApproximationTae Jun Ham, Sungjun Jung, Seonghak Kim, Young H. Oh 等HPCA 2020 · 被引用 241 次
- PREMA: A Predictive Multi-Task Scheduling Algorithm For Preemptible Neural Processing UnitsYujeong Choi, Minsoo RhuHPCA 2020 · 被引用 150 次
- A Multi-Neural Network Acceleration ArchitectureEunjin Baek, Dongup Kwon, Jangwoo KimISCA 2020 · 被引用 110 次
相关 Paper
- USHER: Holistic Interference Avoidance for Resource Optimized ML InferenceSudipta Saha Shubha, Haiying Shen, Anand P. IyerOSDI 2024 · 被引用 35 次
- LaLaRAND: Flexible Layer-by-Layer CPU/GPU Scheduling for Real-Time DNN TasksWoosung Kang, Kilho Lee, Jinkyu Lee, Insik Shin 等RTSS 2021 · 被引用 68 次
- BOER: Enhancing Resource Utilization for Deep Learning Inference with Hybrid Spatial GPU SharingBowen Zhang, Yuhang Wang, Zhuozhao LiSC 2025 · 被引用 4 次
- Automated End-to-End Model Serving with Cooperative Compilation and SchedulingYikang Zhang, Junlong Chen, Wei Wang, Jia Liu 等EuroSys 2026
- EDA: Energy-Efficient Inter-Layer Model Compilation for Edge DNN Inference AccelerationBo Ren Pao, I-Chia Chen, En-Hao Chang, Tsung Tai YehHPCA 2025 · 被引用 1 次
