Optimus: Warming Serverless ML Inference via Inter-Function Model Transformation
Zicong Hong, Jian Lin, Song Guo, Sifu Luo, Wuhui Chen, Roger Wattenhofer, Yue Yu
Abstract
Serverless ML inference is an emerging cloud computing paradigm for low-cost, easy-to-manage inference services. In serverless ML inference, each call is executed in a container; however, the cold start of containers results in long inference delays. Unfortunately, most existing works do not work well because they still need to load models into containers from scratch, which is the bottleneck based on our observations. Therefore, this paper proposes a low-latency serverless ML inference system called Optimus via a new container management scheme. Our key insight is that the loading of a new model can be significantly accelerated when using an existing model with a similar structure in a warm but idle container. We thus develop a novel idea of inter-function model transformation for serverless ML inference, which delves into models within containers at a finer granularity of operations, designs a set of in-container meta-operators for both CNN and transformer model transformation, and develops an efficient scheduling algorithm with linear complexity for a low-cost transformation strategy. Our evaluations on thousands of models show that Optimus reduces inference latency by 24.00% 47.56% in both simulated and real-world workloads compared to state-of-the-art work.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 27bd3a8b-7cc4-49b1-a020-04ad69c5f8eeCited by top-tier papers6
- Improving GPU Sharing Performance through Adaptive Bubbleless Spatial-Temporal SharingShulai Zhang, Quan Chen, Weihao Cui, Han Zhao et al.EuroSys 2025 · 19 citations
- Medusa: Accelerating Serverless LLM Inference with MaterializationShaoxun Zeng, Minhui Xie, Shiwei Gao, Youmin Chen et al.ASPLOS 2025 · 12 citations
- Making Serverless Pay-For-Use a Reality with LeopardTingjia Cao, Andrea C. Arpaci-Dusseau, Remzi H. Arpaci-Dusseau, Tyler Caraza-HarterNSDI 2025 · 11 citations
- Online Container Caching with Late-Warm for IoT Data ProcessingGuopeng Li, Haisheng Tan, Xuan Zhang, Chi Zhang et al.ICDE 2024 · 4 citations
- Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache ManagementQianli Liu, Zicong Hong, Peng Li, Fahao Chen et al.INFOCOM 2025 · 4 citations
Builds on18
- Serverless in the Wild: Characterizing and Optimizing the Serverless Workload at a Large Cloud ProviderMohammad Shahrad, Rodrigo Fonseca, Iñigo Goiri, Gohar Irfan Chaudhry et al.USENIX ATC 2020 · 946 citations
- NAS-Bench-201: Extending the Scope of Reproducible Neural Architecture SearchXuanyi Dong, Yi YangICLR 2020 · 825 citations
- Catalyzer: Sub-millisecond Startup for Serverless Computing with Initialization-less BootingDong Du, Tianyi Yu, Yubin Xia, Binyu Zang et al.ASPLOS 2020 · 280 citations
- FaasCache: keeping serverless computing alive with greedy-dual cachingAlexander Fuerst, Prateek SharmaASPLOS 2021 · 223 citations
- Batch: machine learning inference serving on serverless platforms with adaptive batchingAhsan Ali, Riccardo Pinciroli, Feng Yan, Evgenia SmirniSC 2020 · 184 citations
Related papers
- Help Rather Than Recycle: Alleviating Cold Startup in Serverless Computing Through Inter-Function Container SharingZijun Li, Linsong Guo, Quan Chen, Jiagan Cheng et al.USENIX ATC 2022 · 135 citations
- Rocket: Warming Serverless Inference via Hierarchical ML Artifact Pre-loading and SharingXiaofei Yue, Song Yang, Fan Li, Youqi Li et al.INFOCOM 2026 · 2 citations
- ServerlessLLM: Low-Latency Serverless Inference for Large Language ModelsYao Fu, Leyang Xue, Yeqi Huang, Andrei-Octavian Brabete et al.OSDI 2024 · 125 citations
- FaSei: Fast Serverless Edge Inference with Synergistic Lazy Loading and Layer-wise CachingZhaowu Huang, Fang Dong, Xiaolin Guo, Daheng YinINFOCOM 2025 · 4 citations
- HydraServe: Minimizing Cold Start Latency for Serverless LLM Serving in Public CloudsChiheng Lou, Sheng Qi, Chao Jin, Dapeng Nie et al.NSDI 2026 · 22 citations
