USENIX ATC2024顶会
StreamBox: A Lightweight GPU SandBox for Serverless Inference Workflow
Hao Wu, Yue Yu, Junxiao Deng, Shadi Ibrahim, Song Wu, Hao Fan, Ziyue Cheng, Hai Jin
摘要
The dynamic workload and latency sensitivity of DNN inference drive a trend toward exploiting serverless computing for scalable DNN inference serving. Usually, GPUs are spatially partitioned to serve multiple co-located functions. However, existing serverless inference systems isolate functions in separate monolithic GPU runtimes (e.g., CUDA context), which is too heavy for short-lived and fine-grained functions, leading to a high startup latency, a large memory footprint, and expensive inter-function communication. In this paper, we present StreamBox, a new lightweight GPU sandbox for serverless inference workflow. StreamBox unleashes the potential of streams and efficiently realizes them for serverless inference by implementing fine-grain and auto-scaling memory management, allowing transparent and efficient intra-GPU communication across functions, and enabling PCIe bandwidth sharing among concurrent streams. Our evaluations over real-world workloads show that StreamBox reduces the GPU memory footprint by up to 82% and improves throughput by 6.7X compared to state-of-the-art serverless inference systems.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Katz: Efficient Workflow Serving for Diffusion Models with Many AdaptersSuyi Li, Lingyun Yang, Xiaoxiao Jiang, Hanfeng Lu 等USENIX ATC 2025 · 被引用 14 次
- SGDRC: Software-Defined Dynamic Resource Control for Concurrent DNN Inference on NVIDIA GPUsYongkang Zhang, Haoxuan Yu, Chenxia Han, Cheng Wang 等PPoPP 2025 · 被引用 9 次
- FluidFaaS: A Dynamic Pipelined Solution for Serverless Computing with Strong Isolation-based GPU SharingXinning Hui, Yuanchao Xu, Xipeng ShenHPDC 2025 · 被引用 2 次
- Efficient Data Passing for Serverless Inference Workflows: A GPU-Centric ApproachHao Wu, Yaochen Liu, Minchen Yu, Qizhen Weng 等EuroSys 2026
- Torpor: GPU-Enabled Serverless Computing for Low-Latency, Resource-Efficient InferenceMinchen Yu, Ao Wang, Dong Chen, Haoxuan Yu 等USENIX ATC 2025
它引用的顶会 Paper26
- Serverless in the Wild: Characterizing and Optimizing the Serverless Workload at a Large Cloud ProviderMohammad Shahrad, Rodrigo Fonseca, Iñigo Goiri, Gohar Irfan Chaudhry 等USENIX ATC 2020 · 被引用 946 次
- Faasm: Lightweight Isolation for Efficient Stateful Serverless ComputingSimon Shillaker, Peter R. PietzuchUSENIX ATC 2020 · 被引用 382 次
- Nightcore: efficient and scalable serverless computing for latency-sensitive, interactive microservicesZhipeng Jia, Emmett WitchelASPLOS 2021 · 被引用 218 次
- Serving Heterogeneous Machine Learning Models on Multi-GPU Servers with Spatio-Temporal SharingSeungbeom Choi, Sunho Lee, Yeonjae Kim, Jongse Park 等USENIX ATC 2022 · 被引用 200 次
- Rammer: Enabling Holistic Deep Learning Compiler Optimizations with rTasksLingxiao Ma, Zhiqiang Xie, Zhi Yang, Jilong Xue 等OSDI 2020 · 被引用 192 次
相关 Paper
- FSD-Inference: Fully Serverless Distributed Inference with Scalable Cloud CommunicationJoe Oakley, Hakan FerhatosmanogluICDE 2024 · 被引用 6 次
- Serverless computing on heterogeneous computersDong Du, Qingyuan Liu, Xueqiang Jiang, Yubin Xia 等ASPLOS 2022 · 被引用 68 次
- Dilu: Enabling GPU Resourcing-on-Demand for Serverless DL Serving via Introspective ElasticityCunchi Lv, Xiao Shi, Zhengyu Lei, Jinyue Huang 等ASPLOS 2025 · 被引用 10 次
- TrEnv: Transparently Share Serverless Execution Environments Across Different Functions and NodesJialiang Huang, Mingxing Zhang, Teng Ma, Zheng Liu 等SOSP 2024 · 被引用 14 次
- Optimus: Warming Serverless ML Inference via Inter-Function Model TransformationZicong Hong, Jian Lin, Song Guo, Sifu Luo 等EuroSys 2024 · 被引用 29 次
