Rocket: Warming Serverless Inference via Hierarchical ML Artifact Pre-loading and Sharing
Xiaofei Yue, Song Yang, Fan Li, Youqi Li, Yu Wang
摘要
Serverless computing is a promising method to serve Machine Learning (ML) inference via on-demand functions. Due to the time- and memory-consuming ML library and model (i.e., ML artifact) loading, serverless inference endures notable startup overhead and memory waste issues. In this paper, we advocate for hierarchical ML artifact pre-loading and sharing to balance loading and memory efficiency. Building on this, we propose Rocket, a serverless ML inference system that accelerates function startup while reducing memory waste. Rocket dynamically pre-loads partial, shared ML artifacts, each implying a hierarchy of trade-offs between the loading latency and memory usage. Specifically, with a dual-timescale invocation prediction, Rocket first estimates the pre-loading timing for each function, and then schedules them via a sharing-aware agglomerative clustering to improve ML artifact sharing efficiency. In particular, Rocket learns to make the online hierarchical pre-loading decision for function containers based on a lightweight contextual bandit algorithm. Finally, we implement Rocket and evaluate it with realistic workloads. Experimental results display that Rocket outperforms existing solutions by up to 38.7% on startup latency and up to 43.8% on memory saving.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- Optimus: Warming Serverless ML Inference via Inter-Function Model TransformationZicong Hong, Jian Lin, Song Guo, Sifu Luo 等EuroSys 2024 · 被引用 29 次
- SPES: Towards Optimizing Performance-Resource Trade-Off for Serverless FunctionsCheryl Lee, Zhouruixin Zhu, Tianyi Yang, Yintong Huo 等ICDE 2024 · 被引用 13 次
- ServerlessLLM: Low-Latency Serverless Inference for Large Language ModelsYao Fu, Leyang Xue, Yeqi Huang, Andrei-Octavian Brabete 等OSDI 2024 · 被引用 125 次
- RainbowCake: Mitigating Cold-starts in Serverless with Layer-wise Container Caching and SharingHanfei Yu, Rohan Basu Roy, Christian Fontenot, Devesh Tiwari 等ASPLOS 2024 · 被引用 69 次
- FaSei: Fast Serverless Edge Inference with Synergistic Lazy Loading and Layer-wise CachingZhaowu Huang, Fang Dong, Xiaolin Guo, Daheng YinINFOCOM 2025 · 被引用 4 次
