Semantic Caching for Low-Cost LLM Serving: From Offline Learning to Online Adaptation
Xutong Liu, Baran Atalar, Xiangxiang Dai, Jinhang Zuo, Siwei Wang, John C. S. Lui, Wei Chen, Carlee Joe-Wong
摘要
Large Language Models (LLMs) are revolutionizing how users interact with information systems, yet their high inference cost poses serious scalability and sustainability challenges. Caching inference responses, allowing them to be retrieved without another forward pass through the LLM, has emerged as one possible solution. Traditional exact-match caching, however, overlooks the semantic similarity between queries, leading to unnecessary recomputation. Semantic caching addresses this by retrieving responses based on semantic similarity, but introduces a fundamentally different cache eviction problem: one must account for mismatch costs between incoming queries and cached responses. Moreover, key system parameters, such as query arrival probabilities and serving costs, are often unknown and must be learned over time. Existing semantic caching methods are largely ad-hoc, lacking theoretical foundations and unable to adapt to real-world uncertainty. In this paper, we present a principled, learning-based framework for semantic cache eviction under unknown query and cost distributions. We formulate both offline optimization and online learning variants of the problem based on the combinatorial multi-armed bandit framework, and develop provably efficient algorithms with state-of-the-art guarantees. We also evaluate our framework on a synthetic dataset, showing that our proposed algorithms perform matching or superior performance compared with baselines.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper14
- SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language ModelsGuangxuan Xiao, Ji Lin, Mickaël Seznec, Hao Wu 等ICML 2023 · 被引用 1,493 次
- Fast Inference from Transformers via Speculative DecodingYaniv Leviathan, Matan Kalman, Yossi MatiasICML 2023 · 被引用 1,472 次
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
- ServerlessLLM: Low-Latency Serverless Inference for Large Language ModelsYao Fu, Leyang Xue, Yeqi Huang, Andrei-Octavian Brabete 等OSDI 2024 · 被引用 125 次
- SpotServe: Serving Generative Large Language Models on Preemptible InstancesXupeng Miao, Chunan Shi, Jiangfei Duan, Xiaoli Xi 等ASPLOS 2024 · 被引用 71 次
相关 Paper
- Learned Prefix Caching for Efficient LLM InferenceDongsheng Yang, Austin T. Li, Kai Li, Wyatt LloydNeurIPS 2025 · 被引用 8 次
- SpindleKV: A Novel KV Cache Reduction Method Balancing Both Shallow and Deep LayersZicong Tang, Luohe Shi, Zuchao Li, Baoyuan Qi 等ACL 2025 · 被引用 7 次
- vCache: Verified Semantic Prompt CachingLuis Gaspar Schroeder, Aditya Desai, Alejandro Cuadron, Kyle Chu 等ICLR 2026 · 被引用 15 次
- When Cache Poisoning Meets LLM Systems: Semantic Cache Poisoning and Its CountermeasuresGuanlong Wu, Taojie Wang, Yao Zhang, Zheng Zhang 等NDSS 2026 · 被引用 6 次
- Randomization Boosts KV Caching, Learning Balances Query Load: A Joint PerspectiveFangzhou Wu, Sandeep Silwal, Qiuyi (Richard) ZhangICLR 2026 · 被引用 3 次
