BlitzScale: Fast and Live Large Model Autoscaling with O(1) Host Caching
Dingyan Zhang, Haotian Wang, Yang Liu, Xingda Wei, Yizhou Shan, Rong Chen, Haibo Chen
Abstract
Model autoscaling is the key mechanism for serverless model-as-a-service, but faces a fundamental trade-off between scaling speed and storage/memory usage for caching parameters, and cannot meet frequent multi-host scaling demands. The root cause is a slow, blocking data plane: scaled instances stop while parameters load.
In this paper, we first show that the data plane—loading model checkpoints to accelerators—can be made fast with no or O (1) caching, by loading parameters through the inter-GPU compute network: (1) its speed is comparable to host cache yet underutilized, and (2) scaling multiple instances needs no or O (1) caching via network-optimized multicast. Second, autoscaling can be made live by shifting the scaling abstraction from coarse-grained instance-level to fine-grained layer-level, allowing us to offload layer computation from overloaded instances to scaled ones before parameters fully load.
Under real-world workloads, BlitzScale achieves up to 94 % lower tail latency than the state-of-the-art autoscaling system (ServerlessLLM), and cuts serving GPU time by 49 % versus non-autoscaling systems like DistServe and vLLM at the same SLA. To ease adoption in ecosystems like vLLM and SGLang, we further build BlitzLoad , a lightweight checkpoint engine that brings BlitzScale ’s data plane to existing serving engines with only a few lines of code changes.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fcb840ee-1a01-48d0-8ab1-7e41c7ac66f0Cited by top-tier papers10
- Prism: Cost-Efficient Multi-LLM Serving via GPU Memory BallooningShan Yu, Yifan Qiao, Mingyuan Ma, Yangmin Li et al.OSDI 2026 · 33 citations
- HydraServe: Minimizing Cold Start Latency for Serverless LLM Serving in Public CloudsChiheng Lou, Sheng Qi, Chao Jin, Dapeng Nie et al.NSDI 2026 · 22 citations
- RollArt: Disaggregated Multi-Task Agentic RL Training at ScaleWei Gao, Yuheng Zhao, Tianyuan Wu, Shaopan Xiong et al.OSDI 2026 · 16 citations
- MSCCL++: Rethinking GPU Communication Abstractions for AI InferenceChangho Hwang, Peng Cheng, Roshan Dathathri, Abhinav Jangda et al.ASPLOS 2026 · 5 citations
- Bidaw: Enhancing Key-Value Caching for Interactive LLM Serving via Bidirectional Computation-Storage AwarenessShipeng Hu, Guangyan Zhang, Yuqi Zhou, Yaya Wei et al.FAST 2026 · 3 citations
Builds on23
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
- Serverless in the Wild: Characterizing and Optimizing the Serverless Workload at a Large Cloud ProviderMohammad Shahrad, Rodrigo Fonseca, Iñigo Goiri, Gohar Irfan Chaudhry et al.USENIX ATC 2020 · 946 citations
- Orca: A Distributed Serving System for Transformer-Based Generative ModelsGyeong-In Yu, Joo Seong Jeong, Geon-Woo Kim, Soojeong Kim et al.OSDI 2022 · 690 citations
- Efficient large-scale language model training on GPU clusters using megatron-LMDeepak Narayanan, Mohammad Shoeybi, Jared Casper, Patrick LeGresley et al.SC 2021 · 576 citations
- Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-ServeAmey Agrawal, Nitin Kedia, Ashish Panwar, Jayashree Mohan et al.OSDI 2024 · 537 citations
Related papers
- ServerlessLLM: Low-Latency Serverless Inference for Large Language ModelsYao Fu, Leyang Xue, Yeqi Huang, Andrei-Octavian Brabete et al.OSDI 2024 · 125 citations
- Medusa: Accelerating Serverless LLM Inference with MaterializationShaoxun Zeng, Minhui Xie, Shiwei Gao, Youmin Chen et al.ASPLOS 2025 · 12 citations
- WarmServe: Enabling One-for-Many GPU Prewarming for Multi-LLM ServingChiheng Lou, Sheng Qi, Rui Kang, Yong Zhang et al.ICML 2026 · 3 citations
- KUNSERVE: Parameter-centric Memory Management for Efficient Memory Overloading Handling in LLM ServingRongxin Cheng, Yuxin Lai, Xingda Wei, Rong Chen et al.EuroSys 2026
- INFaaS: Automated Model-less Inference ServingFrancisco Romero, Qian Li, Neeraja J. Yadwadkar, Christos KozyrakisUSENIX ATC 2021 · 325 citations
