BlitzScale: Fast and Live Large Model Autoscaling with O(1) Host Caching
Dingyan Zhang, Haotian Wang, Yang Liu, Xingda Wei, Yizhou Shan, Rong Chen, Haibo Chen
摘要
Model autoscaling is the key mechanism for serverless model-as-a-service, but faces a fundamental trade-off between scaling speed and storage/memory usage for caching parameters, and cannot meet frequent multi-host scaling demands. The root cause is a slow, blocking data plane: scaled instances stop while parameters load.
In this paper, we first show that the data plane—loading model checkpoints to accelerators—can be made fast with no or O (1) caching, by loading parameters through the inter-GPU compute network: (1) its speed is comparable to host cache yet underutilized, and (2) scaling multiple instances needs no or O (1) caching via network-optimized multicast. Second, autoscaling can be made live by shifting the scaling abstraction from coarse-grained instance-level to fine-grained layer-level, allowing us to offload layer computation from overloaded instances to scaled ones before parameters fully load.
Under real-world workloads, BlitzScale achieves up to 94 % lower tail latency than the state-of-the-art autoscaling system (ServerlessLLM), and cuts serving GPU time by 49 % versus non-autoscaling systems like DistServe and vLLM at the same SLA. To ease adoption in ecosystems like vLLM and SGLang, we further build BlitzLoad , a lightweight checkpoint engine that brings BlitzScale ’s data plane to existing serving engines with only a few lines of code changes.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Prism: Cost-Efficient Multi-LLM Serving via GPU Memory BallooningShan Yu, Yifan Qiao, Mingyuan Ma, Yangmin Li 等OSDI 2026 · 被引用 33 次
- HydraServe: Minimizing Cold Start Latency for Serverless LLM Serving in Public CloudsChiheng Lou, Sheng Qi, Chao Jin, Dapeng Nie 等NSDI 2026 · 被引用 22 次
- RollArt: Disaggregated Multi-Task Agentic RL Training at ScaleWei Gao, Yuheng Zhao, Tianyuan Wu, Shaopan Xiong 等OSDI 2026 · 被引用 16 次
- MSCCL++: Rethinking GPU Communication Abstractions for AI InferenceChangho Hwang, Peng Cheng, Roshan Dathathri, Abhinav Jangda 等ASPLOS 2026 · 被引用 5 次
- Bidaw: Enhancing Key-Value Caching for Interactive LLM Serving via Bidirectional Computation-Storage AwarenessShipeng Hu, Guangyan Zhang, Yuqi Zhou, Yaya Wei 等FAST 2026 · 被引用 3 次
它引用的顶会 Paper23
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
- Serverless in the Wild: Characterizing and Optimizing the Serverless Workload at a Large Cloud ProviderMohammad Shahrad, Rodrigo Fonseca, Iñigo Goiri, Gohar Irfan Chaudhry 等USENIX ATC 2020 · 被引用 946 次
- Orca: A Distributed Serving System for Transformer-Based Generative ModelsGyeong-In Yu, Joo Seong Jeong, Geon-Woo Kim, Soojeong Kim 等OSDI 2022 · 被引用 690 次
- Efficient large-scale language model training on GPU clusters using megatron-LMDeepak Narayanan, Mohammad Shoeybi, Jared Casper, Patrick LeGresley 等SC 2021 · 被引用 576 次
- Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-ServeAmey Agrawal, Nitin Kedia, Ashish Panwar, Jayashree Mohan 等OSDI 2024 · 被引用 537 次
相关 Paper
- ServerlessLLM: Low-Latency Serverless Inference for Large Language ModelsYao Fu, Leyang Xue, Yeqi Huang, Andrei-Octavian Brabete 等OSDI 2024 · 被引用 125 次
- Medusa: Accelerating Serverless LLM Inference with MaterializationShaoxun Zeng, Minhui Xie, Shiwei Gao, Youmin Chen 等ASPLOS 2025 · 被引用 12 次
- WarmServe: Enabling One-for-Many GPU Prewarming for Multi-LLM ServingChiheng Lou, Sheng Qi, Rui Kang, Yong Zhang 等ICML 2026 · 被引用 3 次
- KUNSERVE: Parameter-centric Memory Management for Efficient Memory Overloading Handling in LLM ServingRongxin Cheng, Yuxin Lai, Xingda Wei, Rong Chen 等EuroSys 2026
- INFaaS: Automated Model-less Inference ServingFrancisco Romero, Qian Li, Neeraja J. Yadwadkar, Christos KozyrakisUSENIX ATC 2021 · 被引用 325 次
