Lune

INFOCOM2025顶会

Online Context Caching for Distributed Large Language Models Serving

Bin Gao, Zhuomin He, Yizhen Yao, Zhanzhi Lou Lou, Zhi Zhou, Weng-Fai Wong

2025年份
2被引次数

摘要

Large language models (LLMs) based on transformer architectures have demonstrated exceptional performance across various generative tasks. However, the significant GPU resources required for LLM inference pose financial challenges for large-scale deployment. Context caching has been proposed to enhance cost-efficiency by storing intermediate key and value (KV) pairs in cost-effective storage mediums, which can be reused to accelerate inference when requests share prefixes. While promising in single-instance applications, context caching in distributed LLM serving systems introduces unique challenges. Firstly, context caching decisions across time slots are inter-dependent, affecting overall system efficiency due to potential cache misses. Secondly, request scheduling complexities arise, leading to load balancing issues among instances. We address these challenges by formulating an online optimization problem that jointly decides KV cache placement and request scheduling to minimize inference costs across time slots. Given the NP-hard nature of this problem, we propose a framework leveraging regularization, linear relaxation, and randomized rounding techniques. Our solution achieves a competitive ratio near the offline optimum. Experimental results in a distributed LLM serving system demonstrate significant performance improvements over baseline methods.

问问这篇 Paper

问问你的智能体。

Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。

可以从这些问题问起

智能体调用

Lunesearch_papers

在 Lune 里问

免费开始,无需绑卡

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖