Online Context Caching for Distributed Large Language Models Serving
Bin Gao, Zhuomin He, Yizhen Yao, Zhanzhi Lou Lou, Zhi Zhou, Weng-Fai Wong
Abstract
Large language models (LLMs) based on transformer architectures have demonstrated exceptional performance across various generative tasks. However, the significant GPU resources required for LLM inference pose financial challenges for large-scale deployment. Context caching has been proposed to enhance cost-efficiency by storing intermediate key and value (KV) pairs in cost-effective storage mediums, which can be reused to accelerate inference when requests share prefixes. While promising in single-instance applications, context caching in distributed LLM serving systems introduces unique challenges. Firstly, context caching decisions across time slots are inter-dependent, affecting overall system efficiency due to potential cache misses. Secondly, request scheduling complexities arise, leading to load balancing issues among instances. We address these challenges by formulating an online optimization problem that jointly decides KV cache placement and request scheduling to minimize inference costs across time slots. Given the NP-hard nature of this problem, we propose a framework leveraging regularization, linear relaxation, and randomized rounding techniques. Our solution achieves a competitive ratio near the offline optimum. Experimental results in a distributed LLM serving system demonstrate significant performance improvements over baseline methods.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 8f01bb85-bb5a-4f05-a613-f5823c0ab5ffRelated papers
- Compute or Load KV Cache? Why Not Both?Shuowei Jin, Xueshen Liu, Qingzhao Zhang, Zhuoqing MaoICML 2025
- LLM Query Scheduling with Prefix Reuse and Latency ConstraintsGregory Dexter, Shao Tang, Ata Fatahi Baarzi, Qingquan Song et al.NeurIPS 2025 · 10 citations
- Strata: Hierarchical Context Caching for Long Context Language Model ServingZhiqiang Xie, Ziyi Xu, Mark Zhao, Yuwei An et al.OSDI 2026 · 40 citations
- DualMap: Enabling Both Cache Affinity and Load Balancing for Distributed LLM ServingYing Yuan, Pengfei Zuo, Bo Wang, Zhangyu Chen et al.ICLR 2026 · 10 citations
- Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache ManagementQianli Liu, Zicong Hong, Peng Li, Fahao Chen et al.INFOCOM 2025 · 4 citations
