PREMA: A Predictive Multi-Task Scheduling Algorithm For Preemptible Neural Processing Units
Yujeong Choi, Minsoo Rhu
摘要
To amortize cost, cloud vendors providing DNN acceleration as a service to end-users employ consolidation and virtualization to share the underlying resources among multiple DNN service requests. This paper makes a case for a "preemptible" neural processing unit (NPU) and a "predictive" multi-task scheduler to meet the latency demands of high-priority inference while maintaining high throughput. We evaluate both the mechanisms that enable NPUs to be preemptible and the policies that utilize them to meet scheduling objectives. We show that preemptive NPU multi-tasking can achieve an average 7.8×, 1.4×, and 4.8× improvement in latency, throughput, and SLA satisfaction, respectively.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper36
- Serving Heterogeneous Machine Learning Models on Multi-GPU Servers with Spatio-Temporal SharingSeungbeom Choi, Sunho Lee, Yeonjae Kim, Jongse Park 等USENIX ATC 2022 · 被引用 200 次
- DeepRecSys: A System for Optimizing End-To-End At-Scale Neural Recommendation InferenceUdit Gupta, Samuel Hsia, Vikram Saraph, Xiaodong Wang 等ISCA 2020 · 被引用 149 次
- Heterogeneous Dataflow Accelerators for Multi-DNN WorkloadsHyoukjun Kwon, Liangzhen Lai, Michael Pellauer, Tushar Krishna 等HPCA 2021 · 被引用 143 次
- Planaria: Dynamic Architecture Fission for Spatial Multi-Tenant Acceleration of Deep Neural NetworksSoroush Ghodrati, Byung Hoon Ahn, Joon Kyung Kim, Sean Kinzer 等MICRO 2020 · 被引用 120 次
- Centaur: A Chiplet-based, Hybrid Sparse-Dense Accelerator for Personalized RecommendationsRanggi Hwang, Taehun Kim, Youngeun Kwon, Minsoo RhuISCA 2020 · 被引用 94 次
相关 Paper
- Harpagon: Minimizing DNN Serving Cost via Efficient Dispatching, Scheduling and SplittingZhixin Zhao, Yitao Hu, Ziqi Gong, Guotao Yang 等INFOCOM 2025 · 被引用 2 次
- Hardware-Assisted Virtualization of Neural Processing Units for Cloud PlatformsYuqi Xue, Yiqi Liu, Lifeng Nai, Jian HuangMICRO 2024 · 被引用 10 次
- Layerweaver: Maximizing Resource Utilization of Neural Processing Units via Layer-Wise SchedulingYoung H. Oh, Seonghak Kim, Yunho Jin, Sam Son 等HPCA 2021 · 被引用 46 次
- Microsecond-scale Preemption for Concurrent GPU-accelerated DNN InferencesMingcong Han, Hanze Zhang, Rong Chen, Haibo ChenOSDI 2022 · 被引用 153 次
- Lazy Batching: An SLA-aware Batching System for Cloud Machine Learning InferenceYujeong Choi, Yunseong Kim, Minsoo RhuHPCA 2021 · 被引用 65 次
