HotPrefix: Hotness-Aware KV Cache Scheduling for Efficient Prefix Sharing in LLM Inference Systems
Yuhang Li, Rong Gu, Chengying Huan, Zhibin Wang, Renjie Yao, Chen Tian, Guihai Chen
Abstract
Prompt engineering techniques are widely used to enhance the generation quality of large language models (LLMs). However, the long prompts significantly increase inference latency and reduce inference throughput. Since many prompts share common prefixes, prefix sharing has been proposed to reuse shared prefix KV caches during inference. Nevertheless, the large number of prefix KV caches and the limited GPU memory capacity make it impractical to store all prefix KV caches in GPU memory. This limitation necessitates the use of external memory storage strategies, which often suffer from high I/O overhead and frequent cache misses with traditional approaches.
To address these challenges, this paper proposes HotPrefix, a hotness-aware KV cache scheduling framework designed for efficient prefix sharing in LLM inference systems. HotPrefix introduces three core innovations: (1) Dynamic Hotness Tracking, which dynamically monitors and updates the hotness of prefix tree nodes over time; (2) Selective KV Cache Admission, which evaluates evicted KV caches from GPU memory, retaining only high-hotness caches in CPU memory to expand GPU memory capacity and reduce KV cache transfer overhead; (3) Hotness Promotion, which periodically promotes high-hotness prefix tree KV caches from CPU memory to GPU memory. This is combined with an efficient pipeline strategy for I/O and computation, ensuring GPU memory is allocated to the most critical prefixes while masking the I/O overhead associated with KV cache transmission. These mechanisms significantly improve cache hit rates, reduce inference latency, and enhance throughput. Implemented in the SGLang framework, HotPrefix reduces inference latency and increases throughput by up to 2.25× compared with vLLM with prefix sharing enabled. Against SGLang, it achieves up to 2× latency reduction and throughput increase.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 338f3988-09db-42cb-958d-bdd72e8d4265Cited by top-tier papers4
- CoDec: Prefix-Shared Decoding Kernel for LLMsZhibin Wang, Rui Ning, Chao Fang, Zhonghui Zhang et al.SIGMOD 2026 · 8 citations
- KVServe: Service-Aware KV Cache Compression for Communication-Efficient Disaggregated LLM ServingZedong Liu, Xinyang Ma, Dejun Luo, Hairui Zhao et al.SIGCOMM 2026 · 7 citations
- AlignedServe: Orchestrating Prefix-aware Batching to Build a High-throughput and Computing-efficient LLM Serving SystemFengyao Bai, Hongbin Zhang, Zhitao Chen, Jiangsu Du et al.SIGMOD 2026 · 3 citations
- Efficient Cooperation-Aware Key and Value Management for LLM InferenceQiheng Sun, Hongwei Zhang, Junxu Liu, Haocheng Xia et al.VLDB 2026
Builds on29
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran et al.NeurIPS 2023 · 5,068 citations
Related papers
- PKAS: Predictive KVCache-Aware Scheduling for Faster LLM and Transformer InferencesJie Ye, Avinash Maurya, Krishna Teja Chitty-Venkata, Bogdan Nicolae et al.HPDC 2026
- ChunkAttention: Efficient Self-Attention with Prefix-Aware KV Cache and Two-Phase PartitionLu Ye, Ze Tao, Yong Huang, Yang LiACL 2024
- LLM Query Scheduling with Prefix Reuse and Latency ConstraintsGregory Dexter, Shao Tang, Ata Fatahi Baarzi, Qingquan Song et al.NeurIPS 2025 · 10 citations
- PAT: Accelerating LLM Decoding via Prefix-Aware Attention with Resource Efficient Multi-Tile KernelJinjun Yi, Zhixin Zhao, Yitao Hu, Ke Yan et al.ASPLOS 2026 · 1 citation
- KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent WorkflowsZaifeng Pan, Ajjkumar Patel, Yipeng Shen, Zhengding Hu et al.NeurIPS 2025 · 77 citations
