KUNSERVE: Parameter-centric Memory Management for Efficient Memory Overloading Handling in LLM Serving
Rongxin Cheng, Yuxin Lai, Xingda Wei, Rong Chen, Haibo Chen
Abstract
Serving LLMs with a cluster of GPUs is common nowadays, where the serving system must meet strict latency SLOs required by applications. However, the stateful nature of LLM serving requires maintaining huge states (i.e., KVCache) in limited GPU memory. Under spikes in real-world workloads, GPU memory can be easily overloaded, leading to orders of magnitude higher response latency due to queuing introduced by waiting for KVCache to be reclaimed. Prior KVCachecentric approaches handle overloading by dropping, migrating, or swapping KVCache. These methods fail to release sufficient memory quickly with requests still queued.
This paper proposes the first parameter-centric approach to handling overloading by selectively dropping replicated parameters to instantly free memory for requests, based on an unnoticed observation that model parameters are commonly replicated across GPUs for serving LLMs. With additional memory, all requests can be served with a larger batch without queuing. To make the parameter-centric approach correct and efficient, we cooperatively execute requests on GPUs with a complete copy of parameters using pipeline parallelism, and derive an appropriate drop plan without unnecessary cooperation. We also design techniques to minimize the performance overhead due to pipeline parallelism with the execution patterns of requests under drop. Evaluations show that KUN-SERVE reduces the tail TTFT of requests under overloading by up to 72.2 × compared to the state-of-the-art systems including Llumnix, vLLM and InferCept.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 99281e73-4dbb-4599-a3c1-7f366a8670cdCited by top-tier papers2
- WarmServe: Enabling One-for-Many GPU Prewarming for Multi-LLM ServingChiheng Lou, Sheng Qi, Rui Kang, Yong Zhang et al.ICML 2026 · 3 citations
- Bullet: Boosting GPU Utilization for LLM Serving via Dynamic Spatial-Temporal OrchestrationZejia Lin, Hongxin Xu, Guanyi Chen, Zhiguang Chen et al.ASPLOS 2026 · 2 citations
Builds on14
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 2,600 citations
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
- Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-ServeAmey Agrawal, Nitin Kedia, Ashish Panwar, Jayashree Mohan et al.OSDI 2024 · 537 citations
- AlpaServe: Statistical Multiplexing with Model Parallelism for Deep Learning ServingZhuohan Li, Lianmin Zheng, Yinmin Zhong, Vincent Liu et al.OSDI 2023 · 211 citations
Related papers
- ConServe: Fine-Grained GPU Harvesting for LLM Online and Offline Co-ServingYifan Qiao, Shan Yu, Shu Anzai, Haoran Ma et al.ICML 2026 · 7 citations
- Serving Hybrid LLM Loads with SLO Guarantees Using CPU-GPU Attention PiggybackingZizhao Mo, Junlin Chen, Huanle Xu, ChengZhong XuSIGMOD 2026 · 1 citation
- Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache ManagementQianli Liu, Zicong Hong, Peng Li, Fahao Chen et al.INFOCOM 2025 · 4 citations
- PKAS: Predictive KVCache-Aware Scheduling for Faster LLM and Transformer InferencesJie Ye, Avinash Maurya, Krishna Teja Chitty-Venkata, Bogdan Nicolae et al.HPDC 2026
- High Throughput and Low Latency LLM Serving via Adaptive KV CachingWenyan Chen, Chengzhi Lu, Huanle Xu, Kejiang Ye et al.EuroSys 2026
