FluidFaaS: A Dynamic Pipelined Solution for Serverless Computing with Strong Isolation-based GPU Sharing
Xinning Hui, Yuanchao Xu, Xipeng Shen
Abstract
Prompted by the rise of artificial intelligence (AI) or machine learning (ML), more serverless workloads demand efficient GPU support. Recent years have witnessed a shift of interest from weak isolation-based methods, such as Multi-Process Service (MPS), to strong isolation-based methods, such as Multi-Instance GPU (MIG), for GPU support on serverless platforms, thanks to concerns about performance interference and security. The current MIG-based solution for serverless computing is, however, subject to severe GPU resource fragmentation and under-utilization. This paper identifies the reason as the gap between current MIG supports in serverless computing and the rigid constraints in MIG (re)configurations. It proposes FluidFaaS, a solution that enables flexible MIG management for serverless computing. Through a novel programming system support, FluidFaaS enables fine-grained resource assignment to the components within a serverless function, based on which, it equips the invokers with runtime support that constructs pipelines on MIGs on the fly for a serverless function. The innovations, along with a hotness-aware eviction-based time sharing of MIG slices, significantly reduce GPU resource fragmentation and enhance system throughput. Evaluations demonstrate that FluidFaaS outperforms the state-of-the-art solutions by 25%-75% in throughput while achieving up to 90% higher SLO hit rates in various workloads.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cf0bf59c-44c2-4b3a-a5c9-ad01a2c603cbCited by top-tier papers1
Ask how each one uses itBuilds on17
- Serverless in the Wild: Characterizing and Optimizing the Serverless Workload at a Large Cloud ProviderMohammad Shahrad, Rodrigo Fonseca, Iñigo Goiri, Gohar Irfan Chaudhry et al.USENIX ATC 2020 · 946 citations
- Serving DNNs like Clockwork: Performance Predictability from the Bottom UpArpan Gujarati, Reza Karimi, Safya Alzayat, Wei Hao et al.OSDI 2020 · 392 citations
- AntMan: Dynamic Scaling on GPU Clusters for Deep LearningWencong Xiao, Shiru Ren, Yong Li, Yang Zhang et al.OSDI 2020 · 260 citations
- Nightcore: efficient and scalable serverless computing for latency-sensitive, interactive microservicesZhipeng Jia, Emmett WitchelASPLOS 2021 · 218 citations
- SONIC: Application-aware Data Passing for Chained Serverless ApplicationsAshraf Mahgoub, Karthick Shankar, Subrata Mitra, Ana Klimovic et al.USENIX ATC 2021 · 170 citations
Related papers
- gShare: Efficient GPU Sharing with Aggressive Scheduling in Multi-tenant FaaS platformYanan Yang, Zhengxiong Jiang, Meiqi Zhu, Hongqiang Xu et al.ASPLOS 2026 · 1 citation
- Dilu: Enabling GPU Resourcing-on-Demand for Serverless DL Serving via Introspective ElasticityCunchi Lv, Xiao Shi, Zhengyu Lei, Jinyue Huang et al.ASPLOS 2025 · 10 citations
- FlexPipe: Adapting Dynamic LLM Serving Through Inflight Pipeline Refactoring in Fragmented Serverless ClustersYanying Lin, Shijie Peng, Chengzhi Lu, ChengZhong Xu et al.EuroSys 2026 · 4 citations
- ParvaGPU: Efficient Spatial GPU Sharing for Large-Scale DNN Inference in Cloud EnvironmentsMunkyu Lee, Sihoon Seong, Minki Kang, Jihyuk Lee et al.SC 2024 · 19 citations
- ESG: Pipeline-Conscious Efficient Scheduling of DNN Workflows on Serverless Platforms with Shareable GPUsXinning Hui, Yuanchao Xu, Zhishan Guo, Xipeng ShenHPDC 2024 · 11 citations
