HSM: A Hybrid Slowdown Model for Multitasking GPUs
Xia Zhao, Magnus Jahre, Lieven Eeckhout
Abstract
Graphics Processing Units (GPUs) are increasingly widely used in the cloud to accelerate compute-heavy tasks. However, GPU-compute applications stress the GPU architecture in different ways -leading to suboptimal resource utilization when a single GPU is used to run a single application. One solution is to use the GPU in a multitasking fashion to improve utilization. Unfortunately, multitasking leads to destructive interference between co-running applications which causes fairness issues and Quality-of-Service (QoS) violations.
We propose the Hybrid Slowdown Model (HSM) to dynamically and accurately predict application slowdown due to interference. HSM overcomes the low accuracy of prior white-box models, and training and implementation overheads of pure black-box models, with a hybrid approach. More specifically, the white-box component of HSM builds upon the fundamental insight that effective bandwidth utilization is proportional to DRAM row buffer hit rate, and the black-box component of HSM uses linear regression to relate row buffer hit rate to performance. HSM accurately predicts application slowdown with an average error of 6.8%, a significant improvement over the current state-of-the-art. In addition, we use HSM to guide various resource management schemes in multitasking GPUs: HSM-Fair significantly improves fairness (by 1.59× on average) compared to even partitioning, whereas HSM-QoS improves system throughput (by 18.9% on average) compared to proportional SM partitioning while maintaining the QoS target for the high-priority application in challenging mixed memory/compute-bound multi-program workloads.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d37a7c2a-44cd-47a9-9257-b3cc0f4cdf7bCited by top-tier papers9
- Enable simultaneous DNN services based on deterministic operator overlap and precise latency predictionWeihao Cui, Han Zhao, Quan Chen, Ningxin Zheng et al.SC 2021 · 62 citations
- VELTAIR: towards high-performance multi-tenant deep learning services via adaptive compilation and schedulingZihan Liu, Jingwen Leng, Zhihui Zhang, Quan Chen et al.ASPLOS 2022 · 52 citations
- Improving GPU Sharing Performance through Adaptive Bubbleless Spatial-Temporal SharingShulai Zhang, Quan Chen, Weihao Cui, Han Zhao et al.EuroSys 2025 · 19 citations
- Towards High-Goodput LLM Serving with Prefill-decode MultiplexingYukang Chen, Weihao Cui, Han Zhao, Ziyi Xu et al.ASPLOS 2026 · 18 citations
- PCCS: Processor-Centric Contention-aware Slowdown Model for Heterogeneous System-on-ChipsYuanchao Xu, Mehmet Esat Belviranli, Xipeng Shen, Jeffrey S. VetterMICRO 2021 · 15 citations
Related papers
- UGPU: Dynamically Constructing Unbalanced GPUs for Enhanced Resource EfficiencyXia Zhao, Guangda Zhang, Lu Wang, Huadong DaiISCA 2025 · 1 citation
- GPU Scale-Model SimulationHossein SeyyedAghaei, Mahmood Naderan-Tahan, Lieven EeckhoutHPCA 2024 · 13 citations
- GFS: A Preemption-aware Scheduling Framework for GPU Clusters with Predictive Spot Instance ManagementJiaang Duan, Shenglin Xu, Shiyou Qian, Dingyu Yang et al.ASPLOS 2026 · 1 citation
- Improving GPU Multi-tenancy with Page Walk StealingB Pratheek, Neha Jawalkar, Arkaprava BasuHPCA 2021 · 26 citations
- HyFiSS: A Hybrid Fidelity Stall-Aware Simulator for GPGPUsJianchao Yang, Mei Wen, Dong Chen, Zhaoyun Chen et al.MICRO 2024 · 8 citations
