StreamBox: A Lightweight GPU SandBox for Serverless Inference Workflow
Hao Wu, Yue Yu, Junxiao Deng, Shadi Ibrahim, Song Wu, Hao Fan, Ziyue Cheng, Hai Jin
Abstract
The dynamic workload and latency sensitivity of DNN inference drive a trend toward exploiting serverless computing for scalable DNN inference serving. Usually, GPUs are spatially partitioned to serve multiple co-located functions. However, existing serverless inference systems isolate functions in separate monolithic GPU runtimes (e.g., CUDA context), which is too heavy for short-lived and fine-grained functions, leading to a high startup latency, a large memory footprint, and expensive inter-function communication. In this paper, we present StreamBox, a new lightweight GPU sandbox for serverless inference workflow. StreamBox unleashes the potential of streams and efficiently realizes them for serverless inference by implementing fine-grain and auto-scaling memory management, allowing transparent and efficient intra-GPU communication across functions, and enabling PCIe bandwidth sharing among concurrent streams. Our evaluations over real-world workloads show that StreamBox reduces the GPU memory footprint by up to 82% and improves throughput by 6.7X compared to state-of-the-art serverless inference systems.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8818fad1-642a-41c9-af35-200c3ff40258Cited by top-tier papers6
- Katz: Efficient Workflow Serving for Diffusion Models with Many AdaptersSuyi Li, Lingyun Yang, Xiaoxiao Jiang, Hanfeng Lu et al.USENIX ATC 2025 · 14 citations
- SGDRC: Software-Defined Dynamic Resource Control for Concurrent DNN Inference on NVIDIA GPUsYongkang Zhang, Haoxuan Yu, Chenxia Han, Cheng Wang et al.PPoPP 2025 · 9 citations
- FluidFaaS: A Dynamic Pipelined Solution for Serverless Computing with Strong Isolation-based GPU SharingXinning Hui, Yuanchao Xu, Xipeng ShenHPDC 2025 · 2 citations
- Efficient Data Passing for Serverless Inference Workflows: A GPU-Centric ApproachHao Wu, Yaochen Liu, Minchen Yu, Qizhen Weng et al.EuroSys 2026
- Torpor: GPU-Enabled Serverless Computing for Low-Latency, Resource-Efficient InferenceMinchen Yu, Ao Wang, Dong Chen, Haoxuan Yu et al.USENIX ATC 2025
Builds on26
- Serverless in the Wild: Characterizing and Optimizing the Serverless Workload at a Large Cloud ProviderMohammad Shahrad, Rodrigo Fonseca, Iñigo Goiri, Gohar Irfan Chaudhry et al.USENIX ATC 2020 · 946 citations
- Faasm: Lightweight Isolation for Efficient Stateful Serverless ComputingSimon Shillaker, Peter R. PietzuchUSENIX ATC 2020 · 382 citations
- Nightcore: efficient and scalable serverless computing for latency-sensitive, interactive microservicesZhipeng Jia, Emmett WitchelASPLOS 2021 · 218 citations
- Serving Heterogeneous Machine Learning Models on Multi-GPU Servers with Spatio-Temporal SharingSeungbeom Choi, Sunho Lee, Yeonjae Kim, Jongse Park et al.USENIX ATC 2022 · 200 citations
- Rammer: Enabling Holistic Deep Learning Compiler Optimizations with rTasksLingxiao Ma, Zhiqiang Xie, Zhi Yang, Jilong Xue et al.OSDI 2020 · 192 citations
Related papers
- FSD-Inference: Fully Serverless Distributed Inference with Scalable Cloud CommunicationJoe Oakley, Hakan FerhatosmanogluICDE 2024 · 6 citations
- Serverless computing on heterogeneous computersDong Du, Qingyuan Liu, Xueqiang Jiang, Yubin Xia et al.ASPLOS 2022 · 68 citations
- Dilu: Enabling GPU Resourcing-on-Demand for Serverless DL Serving via Introspective ElasticityCunchi Lv, Xiao Shi, Zhengyu Lei, Jinyue Huang et al.ASPLOS 2025 · 10 citations
- TrEnv: Transparently Share Serverless Execution Environments Across Different Functions and NodesJialiang Huang, Mingxing Zhang, Teng Ma, Zheng Liu et al.SOSP 2024 · 14 citations
- Optimus: Warming Serverless ML Inference via Inter-Function Model TransformationZicong Hong, Jian Lin, Song Guo, Sifu Luo et al.EuroSys 2024 · 29 citations
