gShare: Efficient GPU Sharing with Aggressive Scheduling in Multi-tenant FaaS platform
Yanan Yang, Zhengxiong Jiang, Meiqi Zhu, Hongqiang Xu, Yujun Wang, Liang Li, Jiansong Zhang, Jie Wu
摘要
Serving ML models with serverless computing has become increasingly popular in recent years. Many of today's cloud vendors have provided GPU functions to meet the performance requirements of different ML scenarios. However, existing production FaaS platforms suffer from GPU under-utilization and high cloud costs due to poor GPU resource management. In this paper, we propose gShare, an on-demand and efficient GPU function management policy in FaaS platforms. gShare provides a fine-grained GPU virtualization solution for a VM-based multi-tenant FaaS environment. It further decouples the GPU resource from the existing CPU-oriented function management paradigm, enabling flexible GPU sharing across tenants. With a user-transparent vGPU remapping design and aggressive request scheduling policy, gShare can significantly improve the cost-efficiency of GPU functions without causing appreciable function performance degradation. Experimental results show that gShare can reduce GPU usage by 43%–63% compared to the baseline while meeting more than 95% of user latency targets. Compared with the state-of-the-art method, it can also reduce cloud costs by 24%-58% while maintaining better function performance, benefiting both the cloud provider and users.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper1
问问它们各自怎么用它相关 Paper
- FluidFaaS: A Dynamic Pipelined Solution for Serverless Computing with Strong Isolation-based GPU SharingXinning Hui, Yuanchao Xu, Xipeng ShenHPDC 2025 · 被引用 2 次
- Torpor: GPU-Enabled Serverless Computing for Low-Latency, Resource-Efficient InferenceMinchen Yu, Ao Wang, Dong Chen, Haoxuan Yu 等USENIX ATC 2025
- KubeShare: A Framework to Manage GPUs as First-Class and Shared Resources in Container CloudTing-An Yeh, Hung-Hsin Chen, Jerry ChouHPDC 2020 · 被引用 56 次
- Towards Resource-Efficient Serverless LLM Inference with SLINFERChuhao Xu, Zijun Li, Quan Chen, Han Zhao 等HPCA 2026
- Dilu: Enabling GPU Resourcing-on-Demand for Serverless DL Serving via Introspective ElasticityCunchi Lv, Xiao Shi, Zhengyu Lei, Jinyue Huang 等ASPLOS 2025 · 被引用 10 次
