HSM: A Hybrid Slowdown Model for Multitasking GPUs
Xia Zhao, Magnus Jahre, Lieven Eeckhout
摘要
Graphics Processing Units (GPUs) are increasingly widely used in the cloud to accelerate compute-heavy tasks. However, GPU-compute applications stress the GPU architecture in different ways -leading to suboptimal resource utilization when a single GPU is used to run a single application. One solution is to use the GPU in a multitasking fashion to improve utilization. Unfortunately, multitasking leads to destructive interference between co-running applications which causes fairness issues and Quality-of-Service (QoS) violations.
We propose the Hybrid Slowdown Model (HSM) to dynamically and accurately predict application slowdown due to interference. HSM overcomes the low accuracy of prior white-box models, and training and implementation overheads of pure black-box models, with a hybrid approach. More specifically, the white-box component of HSM builds upon the fundamental insight that effective bandwidth utilization is proportional to DRAM row buffer hit rate, and the black-box component of HSM uses linear regression to relate row buffer hit rate to performance. HSM accurately predicts application slowdown with an average error of 6.8%, a significant improvement over the current state-of-the-art. In addition, we use HSM to guide various resource management schemes in multitasking GPUs: HSM-Fair significantly improves fairness (by 1.59× on average) compared to even partitioning, whereas HSM-QoS improves system throughput (by 18.9% on average) compared to proportional SM partitioning while maintaining the QoS target for the high-priority application in challenging mixed memory/compute-bound multi-program workloads.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Enable simultaneous DNN services based on deterministic operator overlap and precise latency predictionWeihao Cui, Han Zhao, Quan Chen, Ningxin Zheng 等SC 2021 · 被引用 62 次
- VELTAIR: towards high-performance multi-tenant deep learning services via adaptive compilation and schedulingZihan Liu, Jingwen Leng, Zhihui Zhang, Quan Chen 等ASPLOS 2022 · 被引用 52 次
- Improving GPU Sharing Performance through Adaptive Bubbleless Spatial-Temporal SharingShulai Zhang, Quan Chen, Weihao Cui, Han Zhao 等EuroSys 2025 · 被引用 19 次
- Towards High-Goodput LLM Serving with Prefill-decode MultiplexingYukang Chen, Weihao Cui, Han Zhao, Ziyi Xu 等ASPLOS 2026 · 被引用 18 次
- PCCS: Processor-Centric Contention-aware Slowdown Model for Heterogeneous System-on-ChipsYuanchao Xu, Mehmet Esat Belviranli, Xipeng Shen, Jeffrey S. VetterMICRO 2021 · 被引用 15 次
相关 Paper
- UGPU: Dynamically Constructing Unbalanced GPUs for Enhanced Resource EfficiencyXia Zhao, Guangda Zhang, Lu Wang, Huadong DaiISCA 2025 · 被引用 1 次
- GPU Scale-Model SimulationHossein SeyyedAghaei, Mahmood Naderan-Tahan, Lieven EeckhoutHPCA 2024 · 被引用 13 次
- GFS: A Preemption-aware Scheduling Framework for GPU Clusters with Predictive Spot Instance ManagementJiaang Duan, Shenglin Xu, Shiyou Qian, Dingyu Yang 等ASPLOS 2026 · 被引用 1 次
- Improving GPU Multi-tenancy with Page Walk StealingB Pratheek, Neha Jawalkar, Arkaprava BasuHPCA 2021 · 被引用 26 次
- HyFiSS: A Hybrid Fidelity Stall-Aware Simulator for GPGPUsJianchao Yang, Mei Wen, Dong Chen, Zhaoyun Chen 等MICRO 2024 · 被引用 8 次
