Efficient Data Passing for Serverless Inference Workflows: A GPU-Centric Approach
Hao Wu, Yaochen Liu, Minchen Yu, Qizhen Weng, Junxiao Deng, Yue Yu, Hao Fan, Song Wu, Wei Wang, Hai Jin
Abstract
Serverless computing offers a compelling paradigm for deploying machine learning inference workflows composed of heterogeneous CPU and GPU functions. However, existing data-passing solutions in serverless systems primarily rely on host memory for data exchange (host-centric), leading to substantial data movement and salient I/O overhead. Moreover, modern GPU communication libraries (e.g., NCCL, NVSHMEM, UCX) are ill-suited to serverless environments, suffering from redundant data copies, underutilized transfer bandwidth, and inefficient temporary GPU storage.
In this paper, we present GRouter, a GPU-centric data plane system designed for serverless inference workflows. GRouter first introduces a unified data passing framework that abstracts host-to-GPU and GPU-to-GPU communication while leveraging function placement to reduce redundant copies. It then aggregates available bandwidth across PCIe links, NVLinks, and NICs to enable parallel transfers with performance isolation between functions. GRouter also implements elastic GPU storage that adapts to idle memory availability and varying data transfer demands. Evaluations on real-world inference services show that GRouter reduces data passing latency by up to 87% and improves throughput by up to 1.74× compared to state-of-the-art GPU communication libraries.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 92ccb74d-626d-4d85-a48a-8773f9792c7dBuilds on35
- Serverless in the Wild: Characterizing and Optimizing the Serverless Workload at a Large Cloud ProviderMohammad Shahrad, Rodrigo Fonseca, Iñigo Goiri, Gohar Irfan Chaudhry et al.USENIX ATC 2020 · 946 citations
- Mooncake: Trading More Storage for Less Computation - A KVCache-centric Architecture for Serving LLM ChatbotRuoyu Qin, Zheming Li, Weiran He, Jialei Cui et al.FAST 2025 · 337 citations
- Nightcore: efficient and scalable serverless computing for latency-sensitive, interactive microservicesZhipeng Jia, Emmett WitchelASPLOS 2021 · 218 citations
- Serving Heterogeneous Machine Learning Models on Multi-GPU Servers with Spatio-Temporal SharingSeungbeom Choi, Sunho Lee, Yeonjae Kim, Jongse Park et al.USENIX ATC 2022 · 200 citations
- Batch: machine learning inference serving on serverless platforms with adaptive batchingAhsan Ali, Riccardo Pinciroli, Feng Yan, Evgenia SmirniSC 2020 · 184 citations
Related papers
- StreamBox: A Lightweight GPU SandBox for Serverless Inference WorkflowHao Wu, Yue Yu, Junxiao Deng, Shadi Ibrahim et al.USENIX ATC 2024 · 21 citations
- FSD-Inference: Fully Serverless Distributed Inference with Scalable Cloud CommunicationJoe Oakley, Hakan FerhatosmanogluICDE 2024 · 6 citations
- Torpor: GPU-Enabled Serverless Computing for Low-Latency, Resource-Efficient InferenceMinchen Yu, Ao Wang, Dong Chen, Haoxuan Yu et al.USENIX ATC 2025
- FaaSBoard: Efficient Graph Processing with a Disaggregated Architecture on Serverless ServicesYushi Liu, Yikang Ruan, Letian Ruan, Zijun Li et al.SIGMOD 2026
- <u>G</u>PU <u>i</u>nitiated <u>O</u>penSHMEM: correct and efficient intra-kernel networking for dGPUsKhaled Hamidouche, Michael LeBeanePPoPP 2020 · 17 citations
