Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management
Qianli Liu, Zicong Hong, Peng Li, Fahao Chen, Song Guo
Abstract
Serving large language models (LLMs) for massive users is challenged by the significant memory footprint of the transient state, known as the key-value (KV) cache, which scales with sequence length and number of requests. Instead of renting or buying more expensive GPUs, the load imbalance of the KV cache across GPUs, coupled with recent advances in inter-GPU communication, provides an opportunity to serve more requests via request migration. However, high migration overhead and unpredictable request patterns make it challenging. Therefore, this paper proposes Mell, a memory-efficient LLM serving system via multi-GPU KV cache management. It saves the number of GPUs needed in the system by considering the dynamic KV cache load and the costly request migration. Specifically, we first develop an adaptive request migration mechanism to balance the computational and communication overheads and adapt to diverse resource conditions. Then, we design an online algorithm tailored to a multi-LLM request and multi-GPU scheduling problem with migration enabled. It aims to minimise the required GPUs while limiting the number of migrations. Finally, we implement a prototype of Mell and demonstrate that it reduces the number of GPUs by 31% and increases the GPU utilization by 43% at most compared to existing LLM serving systems.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ef241923-3519-4e65-a5d3-aee58a672d52Cited by top-tier papers4
- Prism: Cost-Efficient Multi-LLM Serving via GPU Memory BallooningShan Yu, Yifan Qiao, Mingyuan Ma, Yangmin Li et al.OSDI 2026 · 33 citations
- PPAI: Enabling Personalized LLM Agent Interoperability for Collaborative Edge IntelligenceZile Wang, Qianli Liu, Kaibin Guo, Haodong Wang et al.INFOCOM 2026 · 4 citations
- MATCH: Modulating Attention via In-Context Retrieval for Long-Context TransformersLinrui Ma, Chun Hei Lo, Xinyu Wang, Peng Lu et al.ACL 2026
- ContrastKV: Robust KV Cache Eviction via Contrastive Signal Fusion for Multi-Query GeneralizationXingchi Chen, Peiyuan Zong, Ziqiang Gao, Qing Li et al.ACL 2026
Builds on26
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han et al.ICLR 2024 · 1,714 citations
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
- H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language ModelsZhenyu Zhang, Ying Sheng, Tianyi Zhou, Tianlong Chen et al.NeurIPS 2023 · 1,003 citations
- Orca: A Distributed Serving System for Transformer-Based Generative ModelsGyeong-In Yu, Joo Seong Jeong, Geon-Woo Kim, Soojeong Kim et al.OSDI 2022 · 690 citations
- FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPUYing Sheng, Lianmin Zheng, Binhang Yuan, Zhuohan Li et al.ICML 2023 · 683 citations
Related papers
- High Throughput and Low Latency LLM Serving via Adaptive KV CachingWenyan Chen, Chengzhi Lu, Huanle Xu, Kejiang Ye et al.EuroSys 2026
- WarmServe: Enabling One-for-Many GPU Prewarming for Multi-LLM ServingChiheng Lou, Sheng Qi, Rui Kang, Yong Zhang et al.ICML 2026 · 3 citations
- DéjàVu: KV-cache Streaming for Fast, Fault-tolerant Generative LLM ServingFoteini Strati, Sara McAllister, Amar Phanishayee, Jakub Tarnawski et al.ICML 2024 · 59 citations
- InfiniGen: Efficient Generative Inference of Large Language Models with Dynamic KV Cache ManagementWonbeom Lee, Jungi Lee, Junghwan Seo, Jaewoong SimOSDI 2024 · 248 citations
- Stateful Large Language Model Serving with PensieveLingfan Yu, Jinkun Lin, Jinyang LiEuroSys 2025 · 23 citations
