V10: Hardware-Assisted NPU Multi-tenancy for Improved Resource Utilization and Fairness
Yuqi Xue, Yiqi Liu, Lifeng Nai, Jian Huang
摘要
Modern cloud platforms have deployed neural processing units (NPUs) like Google Cloud TPUs to accelerate online machine learning (ML) inference services. To improve the resource utilization of NPUs, they allow multiple ML applications to share the same NPU, and developed both time-multiplexed and preemptive-based sharing mechanisms. However, our study with real-world NPUs discloses that these approaches suffer from surprisingly low utilization, due to the lack of support for fine-grained hardware resource sharing in the NPU. Specifically, its separate systolic array and vector unit cannot be fully utilized at the same time, which requires fundamental hardware assistance for supporting multi-tenancy.
In this paper, we present V10, a hardware-assisted NPU multitenancy framework for improving resource utilization, while ensuring fairness for different ML services. We rethink the NPU architecture for supporting multi-tenancy. V10 employs an operator scheduler for enabling concurrent operator executions on the systolic array and the vector unit and offers flexibility for enforcing different priority-based resource-sharing mechanisms. V10 also enables fine-grained operator preemption and lightweight context switch in the NPU. To further improve NPU utilization, V10 also develops a clustering-based workload collocation mechanism for identifying the best-matching ML services on a shared NPU. We implement V10 with an NPU simulator. Our experiments with various ML workloads from MLPerf AI Benchmarks demonstrate that V10 can improve the overall NPU utilization by 1.64×, increase the aggregated throughput by 1.57×, reduce the average latency of ML services by 1.56×, and tail latency by 1.74× on average, in comparison with state-of-the-art NPU multi-tenancy approaches.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- sNPU: Trusted Execution Environments on Integrated NPUsErhu Feng, Dahu Feng, Dong Du, Yubin Xia 等ISCA 2024 · 被引用 13 次
- Topology-Aware Virtualization over Inter-Core Connected Neural Processing UnitsDahu Feng, Erhu Feng, Dong Du, Pinjie Xu 等ISCA 2025 · 被引用 2 次
它引用的顶会 Paper14
- Microsecond-scale Preemption for Concurrent GPU-accelerated DNN InferencesMingcong Han, Hanze Zhang, Rong Chen, Haibo ChenOSDI 2022 · 被引用 153 次
- PREMA: A Predictive Multi-Task Scheduling Algorithm For Preemptible Neural Processing UnitsYujeong Choi, Minsoo RhuHPCA 2020 · 被引用 150 次
- Planaria: Dynamic Architecture Fission for Spatial Multi-Tenant Acceleration of Deep Neural NetworksSoroush Ghodrati, Byung Hoon Ahn, Joon Kyung Kim, Sean Kinzer 等MICRO 2020 · 被引用 120 次
- A Multi-Neural Network Acceleration ArchitectureEunjin Baek, Dongup Kwon, Jangwoo KimISCA 2020 · 被引用 110 次
- Unity: Accelerating DNN Training Through Joint Optimization of Algebraic Transformations and ParallelizationColin Unger, Zhihao Jia, Wei Wu, Sina Lin 等OSDI 2022 · 被引用 105 次
相关 Paper
- Hardware-Assisted Virtualization of Neural Processing Units for Cloud PlatformsYuqi Xue, Yiqi Liu, Lifeng Nai, Jian HuangMICRO 2024 · 被引用 10 次
- Dataflow Mirroring: Architectural Support for Highly Efficient Fine-Grained Spatial Multitasking on Systolic-Array NPUsJounghoo Lee, Jinwoo Choi, Jaeyeon Kim, Jinho Lee 等DAC 2021 · 被引用 39 次
- Serving Heterogeneous Machine Learning Models on Multi-GPU Servers with Spatio-Temporal SharingSeungbeom Choi, Sunho Lee, Yeonjae Kim, Jongse Park 等USENIX ATC 2022 · 被引用 200 次
- CaMDN: Enhancing Cache Efficiency for Multi-tenant DNNs on Integrated NPUsTianhao Cai, Liang Wang, Limin Xiao, Meng Han 等DAC 2025
- MoCA: Memory-Centric, Adaptive Execution for Multi-Tenant Deep Neural NetworksSeah Kim, Hasan Genc, Vadim Vadimovich Nikiforov, Krste Asanovic 等HPCA 2023 · 被引用 36 次
