USENIX ATC2025顶会
KVCache Cache in the Wild: Characterizing and Optimizing KVCache Cache at a Large Cloud Provider
Jiahao Wang, Jinbo Han, Xingda Wei, Sijie Shen, Dingyan Zhang, Chenguang Fang, Rong Chen, Wenyuan Yu, Haibo Chen
摘要
Serving large language models (LLMs) is important for cloud providers, and caching intermediate results (KV$) after processing each request substantially improves serving throughput and latency. However, there is limited understanding of how LLM serving benefits from KV$ caching, where system design decisions like cache eviction policies are highly workload-dependent. In this paper, we present the first systematic characterization of the KV$ workload patterns from one of the leading LLM service providers. We draw observations that were not covered by previous studies focusing on synthetic workloads, including: KV$ reuses are skewed across requests, where reuses between single-turn requests are equally important as multi-turn requests; the reuse time and probability are diverse considering all requests, but for a specific request category, the pattern tends to be predictable; and the overall cache size required for an ideal cache hit ratio is moderate. Based on the characterization, we further propose a workload-aware cache eviction policy that improves the serving performance under real-world traces, especially with limited cache capacity.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- Prism: Cost-Efficient Multi-LLM Serving via GPU Memory BallooningShan Yu, Yifan Qiao, Mingyuan Ma, Yangmin Li 等OSDI 2026 · 被引用 33 次
- BlitzScale: Fast and Live Large Model Autoscaling with O(1) Host CachingDingyan Zhang, Haotian Wang, Yang Liu, Xingda Wei 等OSDI 2025 · 被引用 29 次
- Towards High-Goodput LLM Serving with Prefill-decode MultiplexingYukang Chen, Weihao Cui, Han Zhao, Ziyi Xu 等ASPLOS 2026 · 被引用 18 次
- DualMap: Enabling Both Cache Affinity and Load Balancing for Distributed LLM ServingYing Yuan, Pengfei Zuo, Bo Wang, Zhangyu Chen 等ICLR 2026 · 被引用 10 次
- Cache Me, Catch You: Cache Related Security Threats in LLM Serving FrameworksXiangFan Wu, Lingyun Ying, Guoqiang Chen, Yacong Gu 等NDSS 2026 · 被引用 6 次
它引用的顶会 Paper23
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han 等ICLR 2024 · 被引用 1,714 次
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
- H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language ModelsZhenyu Zhang, Ying Sheng, Tianyi Zhou, Tianlong Chen 等NeurIPS 2023 · 被引用 1,003 次
- Orca: A Distributed Serving System for Transformer-Based Generative ModelsGyeong-In Yu, Joo Seong Jeong, Geon-Woo Kim, Soojeong Kim 等OSDI 2022 · 被引用 690 次
- Efficient large-scale language model training on GPU clusters using megatron-LMDeepak Narayanan, Mohammad Shoeybi, Jared Casper, Patrick LeGresley 等SC 2021 · 被引用 576 次
相关 Paper
- Randomization Boosts KV Caching, Learning Balances Query Load: A Joint PerspectiveFangzhou Wu, Sandeep Silwal, Qiuyi (Richard) ZhangICLR 2026 · 被引用 3 次
- ServeGen: Workload Characterization and Generation of Large Language Model Serving in ProductionYuxing Xiang, Xue Li, Kun Qian, Yan Zhang 等NSDI 2026 · 被引用 58 次
- ConServe: Fine-Grained GPU Harvesting for LLM Online and Offline Co-ServingYifan Qiao, Shan Yu, Shu Anzai, Haoran Ma 等ICML 2026 · 被引用 7 次
- KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent WorkflowsZaifeng Pan, Ajjkumar Patel, Yipeng Shen, Zhengding Hu 等NeurIPS 2025 · 被引用 77 次
- Cost-Efficient Large Language Model Serving for Multi-turn Conversations with CachedAttentionBin Gao, Zhuomin He, Puru Sharma, Qingxuan Kang 等USENIX ATC 2024 · 被引用 273 次
