Medusa: Accelerating Serverless LLM Inference with Materialization
Shaoxun Zeng, Minhui Xie, Shiwei Gao, Youmin Chen, Youyou Lu
摘要
Serverless is a promising paradigm to provide scalable, costefficient, and easy-to-use model inference services. However, the cold start of model inference functions requires loading models to the devices, which incurs high latencies and undermines the benefits of serverless computing. In LLMs, things get even worse since two extra stages are introduced: a KV cache initialization stage that profiles and anticipates memory reservation for KV cache, and a capturing stage which dynamically constructs CUDA graphs for different batch sizes. Both stages are paramount to the inference performance, but become the main culprit of cold start latency.
This paper proposes Medusa to mitigate the long cold start latency through state materialization. Instead of dynamic profiling and construction in the runtime, Medusa materializes the CUDA graphs as well as the information needed by the KV cache initialization in the offline phase, and restores them efficiently in the online phase. Medusa further introduces two novel techniques -offline-online cooperated parameters restoration and triggering-kernels enhanced kernel address restoration -to tackle non-deterministic issues in CUDA graphs. Medusa successfully materializes and restores CUDA graphs across 10 models (with a total of 139364 CUDA graph nodes), and reduces the latency of model loading by 42.5%. Under real-world LLM inference workloads, Medusa reduces the tail latency of the time to first token (TTFT) by 53.0%.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- HydraServe: Minimizing Cold Start Latency for Serverless LLM Serving in Public CloudsChiheng Lou, Sheng Qi, Chao Jin, Dapeng Nie 等NSDI 2026 · 被引用 22 次
- FlexPipe: Adapting Dynamic LLM Serving Through Inflight Pipeline Refactoring in Fragmented Serverless ClustersYanying Lin, Shijie Peng, Chengzhi Lu, ChengZhong Xu 等EuroSys 2026 · 被引用 4 次
- Re-architecting End-host Networking with CXL: Coherence, Memory, and OffloadingHouxiang Ji, Yifan Yuan, Yang Zhou, Ipoom Jeong 等MICRO 2025 · 被引用 3 次
- Towards Resource-Efficient Serverless LLM Inference with SLINFERChuhao Xu, Zijun Li, Quan Chen, Han Zhao 等HPCA 2026
它引用的顶会 Paper23
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
- Serverless in the Wild: Characterizing and Optimizing the Serverless Workload at a Large Cloud ProviderMohammad Shahrad, Rodrigo Fonseca, Iñigo Goiri, Gohar Irfan Chaudhry 等USENIX ATC 2020 · 被引用 946 次
- PyTorch 2: Faster Machine Learning Through Dynamic Python Bytecode Transformation and Graph CompilationJason Ansel, Edward Z. Yang, Horace He, Natalia Gimelshein 等ASPLOS 2024 · 被引用 693 次
- Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-ServeAmey Agrawal, Nitin Kedia, Ashish Panwar, Jayashree Mohan 等OSDI 2024 · 被引用 537 次
- Faasm: Lightweight Isolation for Efficient Stateful Serverless ComputingSimon Shillaker, Peter R. PietzuchUSENIX ATC 2020 · 被引用 382 次
相关 Paper
- Optimus: Warming Serverless ML Inference via Inter-Function Model TransformationZicong Hong, Jian Lin, Song Guo, Sifu Luo 等EuroSys 2024 · 被引用 29 次
- ServerlessLLM: Low-Latency Serverless Inference for Large Language ModelsYao Fu, Leyang Xue, Yeqi Huang, Andrei-Octavian Brabete 等OSDI 2024 · 被引用 125 次
- PASK: Cold Start Mitigation for Inference with Proactive and Selective Kernel Loading on GPUsXuanteng Huang, Jiangsu Du, Nong Xiao, XianWei ZhangDAC 2025
- FaSei: Fast Serverless Edge Inference with Synergistic Lazy Loading and Layer-wise CachingZhaowu Huang, Fang Dong, Xiaolin Guo, Daheng YinINFOCOM 2025 · 被引用 4 次
- SMIless: Serving DAG-based Inference with Dynamic Invocations under Serverless ComputingChengzhi Lu, Huanle Xu, Yudan Li, Wenyan Chen 等SC 2024 · 被引用 8 次
