Medusa: Accelerating Serverless LLM Inference with Materialization
Shaoxun Zeng, Minhui Xie, Shiwei Gao, Youmin Chen, Youyou Lu
Abstract
Serverless is a promising paradigm to provide scalable, costefficient, and easy-to-use model inference services. However, the cold start of model inference functions requires loading models to the devices, which incurs high latencies and undermines the benefits of serverless computing. In LLMs, things get even worse since two extra stages are introduced: a KV cache initialization stage that profiles and anticipates memory reservation for KV cache, and a capturing stage which dynamically constructs CUDA graphs for different batch sizes. Both stages are paramount to the inference performance, but become the main culprit of cold start latency.
This paper proposes Medusa to mitigate the long cold start latency through state materialization. Instead of dynamic profiling and construction in the runtime, Medusa materializes the CUDA graphs as well as the information needed by the KV cache initialization in the offline phase, and restores them efficiently in the online phase. Medusa further introduces two novel techniques -offline-online cooperated parameters restoration and triggering-kernels enhanced kernel address restoration -to tackle non-deterministic issues in CUDA graphs. Medusa successfully materializes and restores CUDA graphs across 10 models (with a total of 139364 CUDA graph nodes), and reduces the latency of model loading by 42.5%. Under real-world LLM inference workloads, Medusa reduces the tail latency of the time to first token (TTFT) by 53.0%.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 775b85f0-3118-43d2-a9f7-6896aba62edaCited by top-tier papers4
- HydraServe: Minimizing Cold Start Latency for Serverless LLM Serving in Public CloudsChiheng Lou, Sheng Qi, Chao Jin, Dapeng Nie et al.NSDI 2026 · 22 citations
- FlexPipe: Adapting Dynamic LLM Serving Through Inflight Pipeline Refactoring in Fragmented Serverless ClustersYanying Lin, Shijie Peng, Chengzhi Lu, ChengZhong Xu et al.EuroSys 2026 · 4 citations
- Re-architecting End-host Networking with CXL: Coherence, Memory, and OffloadingHouxiang Ji, Yifan Yuan, Yang Zhou, Ipoom Jeong et al.MICRO 2025 · 3 citations
- Towards Resource-Efficient Serverless LLM Inference with SLINFERChuhao Xu, Zijun Li, Quan Chen, Han Zhao et al.HPCA 2026
Builds on23
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
- Serverless in the Wild: Characterizing and Optimizing the Serverless Workload at a Large Cloud ProviderMohammad Shahrad, Rodrigo Fonseca, Iñigo Goiri, Gohar Irfan Chaudhry et al.USENIX ATC 2020 · 946 citations
- PyTorch 2: Faster Machine Learning Through Dynamic Python Bytecode Transformation and Graph CompilationJason Ansel, Edward Z. Yang, Horace He, Natalia Gimelshein et al.ASPLOS 2024 · 693 citations
- Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-ServeAmey Agrawal, Nitin Kedia, Ashish Panwar, Jayashree Mohan et al.OSDI 2024 · 537 citations
- Faasm: Lightweight Isolation for Efficient Stateful Serverless ComputingSimon Shillaker, Peter R. PietzuchUSENIX ATC 2020 · 382 citations
Related papers
- Optimus: Warming Serverless ML Inference via Inter-Function Model TransformationZicong Hong, Jian Lin, Song Guo, Sifu Luo et al.EuroSys 2024 · 29 citations
- ServerlessLLM: Low-Latency Serverless Inference for Large Language ModelsYao Fu, Leyang Xue, Yeqi Huang, Andrei-Octavian Brabete et al.OSDI 2024 · 125 citations
- PASK: Cold Start Mitigation for Inference with Proactive and Selective Kernel Loading on GPUsXuanteng Huang, Jiangsu Du, Nong Xiao, XianWei ZhangDAC 2025
- FaSei: Fast Serverless Edge Inference with Synergistic Lazy Loading and Layer-wise CachingZhaowu Huang, Fang Dong, Xiaolin Guo, Daheng YinINFOCOM 2025 · 4 citations
- SMIless: Serving DAG-based Inference with Dynamic Invocations under Serverless ComputingChengzhi Lu, Huanle Xu, Yudan Li, Wenyan Chen et al.SC 2024 · 8 citations
