KRISP: Enabling Kernel-wise RIght-sizing for Spatial Partitioned GPU Inference Servers
Marcus Chow, Ali Jahanshahi, Daniel Wong
摘要
Machine learning (ML) inference workloads present significantly different challenges than ML training workloads. Typically, inference workloads are shorter running and underutilize GPU resources. To overcome this, co-locating multiple instances of a model has been proposed to improve the utilization of GPUs. Co-located models share the GPU through GPU spatial partitioning facilities, such as Nvidia's MPS, MIG, or AMD's CU Masking API. Existing spatially partitioned inference servers create model-wise partitions by "right-sizing" based on a model's latency tolerance to restricting resources. We show that model-wise right-sizing is under-utilized due to varying resource restriction tolerance of individual kernels within an inference pass.
We propose Kernel-wise Right-sizing for Spatial Partitioned GPU Inference Servers (KRISP) to enable kernel-wise right-sizing of spatial partitions at the granularity of individual kernels. We demonstrate that KRISP can support a greater level of concurrently running inference models compared to existing spatially partitioned inference servers. KRISP improves overall throughput by 2x when compared with an isolated inference (1.22x vs prior works) and reduce energy per inference by 33%.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- StreamBox: A Lightweight GPU SandBox for Serverless Inference WorkflowHao Wu, Yue Yu, Junxiao Deng, Shadi Ibrahim 等USENIX ATC 2024 · 被引用 21 次
- Improving GPU Sharing Performance through Adaptive Bubbleless Spatial-Temporal SharingShulai Zhang, Quan Chen, Weihao Cui, Han Zhao 等EuroSys 2025 · 被引用 19 次
- ParvaGPU: Efficient Spatial GPU Sharing for Large-Scale DNN Inference in Cloud EnvironmentsMunkyu Lee, Sihoon Seong, Minki Kang, Jihyuk Lee 等SC 2024 · 被引用 19 次
- throttLL'eM: Predictive GPU Throttling for Energy Efficient LLM Inference ServingAndreas Kosmas Kakolyris, Dimosthenis Masouros, Petros Vavaroutsos, Sotirios Xydis 等HPCA 2025 · 被引用 15 次
- Colocating ML Inference and Training with Fast GPU Memory HandoverJiali Wang, Yankui Wang, Mingcong Han, Rong ChenUSENIX ATC 2025 · 被引用 12 次
它引用的顶会 Paper11
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel 等ICLR 2020 · 被引用 7,418 次
- Serving DNNs like Clockwork: Performance Predictability from the Bottom UpArpan Gujarati, Reza Karimi, Safya Alzayat, Wei Hao 等OSDI 2020 · 被引用 392 次
- Accel-Sim: An Extensible Simulation Framework for Validated GPU ModelingMahmoud Khairy, Zhesheng Shen, Tor M. Aamodt, Timothy G. RogersISCA 2020 · 被引用 366 次
- INFaaS: Automated Model-less Inference ServingFrancisco Romero, Qian Li, Neeraja J. Yadwadkar, Christos KozyrakisUSENIX ATC 2021 · 被引用 325 次
- DeepRecSys: A System for Optimizing End-To-End At-Scale Neural Recommendation InferenceUdit Gupta, Samuel Hsia, Vikram Saraph, Xiaodong Wang 等ISCA 2020 · 被引用 149 次
相关 Paper
- Serving Heterogeneous Machine Learning Models on Multi-GPU Servers with Spatio-Temporal SharingSeungbeom Choi, Sunho Lee, Yeonjae Kim, Jongse Park 等USENIX ATC 2022 · 被引用 200 次
- BOER: Enhancing Resource Utilization for Deep Learning Inference with Hybrid Spatial GPU SharingBowen Zhang, Yuhang Wang, Zhuozhao LiSC 2025 · 被引用 4 次
- Beyond Utilization: Energy-Conscious GPU Sharing for Inference ServingPrasoon Sinha, Dimitrios Liakopoulos, Nathan Lemma, Neeraja J. YadwadkarSOSP 2026
- Splitwise: Efficient Generative LLM Inference Using Phase SplittingPratyush Patel, Esha Choukse, Chaojie Zhang, Aashaka Shah 等ISCA 2024 · 被引用 282 次
- Multiplexing Dynamic Deep Learning Workloads with SLO-awareness in GPU ClustersWenyan Chen, Chengzhi Lu, Huanle Xu, Kejiang Ye 等EuroSys 2025 · 被引用 13 次
