Fast State Restoration in LLM Serving with HCache
Shiwei Gao, Youmin Chen, Jiwu Shu
Abstract
The growing complexity of LLM usage today, e.g., multi-round conversation and retrieval-augmented generation (RAG), makes contextual states (i.e., KV cache) reusable across user requests. Given the capacity constraints of GPU memory, only a limited number of contexts can be cached on GPU for reusing. Existing inference systems typically evict part of the KV cache and restore it by recomputing it from the original tokens or offloading it to host storage for later retrieval, both of which introduce substantial computational or I/O overheads.
We propose HCache, a novel LLM state restoration method. Its key idea is to restore LLM states from intermediate activations and thus utilize computational and I/O resources with low overhead. We enhance HCache with two techniques, including i) a bubble-free restoration scheduler that integrates resource-complementary methods to optimize the balance between computation and IO tasks; and ii) a chunk-based storage manager to address the layout mismatch issue (i.e., layer-before-token saving versus token-before-layer restoration). Our evaluations, conducted using real-world tasks, show that HCache reduces the TTFT by up to 1.93× compared to KV offload while consuming 1.92-2.40× less storage space; compared to token recomputation, HCache achieves up to 5.73× reduction in TTFT.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 97b72281-93bc-4d96-9fec-2c450fb0aa6aCited by top-tier papers9
- RollPacker: Taming Long-Tail Rollouts for RL Post-Training with Tail BatchingWei Gao, Yuheng Zhao, Dakai An, Tianyuan Wu et al.NSDI 2026 · 10 citations
- Bidaw: Enhancing Key-Value Caching for Interactive LLM Serving via Bidirectional Computation-Storage AwarenessShipeng Hu, Guangyan Zhang, Yuqi Zhou, Yaya Wei et al.FAST 2026 · 3 citations
- ORBITFLOW: SLO-Aware Long-Context LLM Serving with Fine-Grained KV Cache ReconfigurationXinyue Ma, Heelim Hong, Taegeon Um, Jongseop Lee et al.VLDB 2026 · 3 citations
- RetroInfer: A Vector Storage Engine for Scalable Long-Context LLM InferenceYaoqi Chen, Jinkai Zhang, Baotong Lu, Qianxi Zhang et al.VLDB 2026 · 1 citation
- Bat: Efficient Generative Recommender Serving with Bipartite AttentionJie Sun, Shaohang Wang, Zimo Zhang, Zhengyu Liu et al.ASPLOS 2026 · 1 citation
Builds on26
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Improving Language Models by Retrieving from Trillions of TokensSebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, Trevor Cai et al.ICML 2022 · 1,629 citations
- SGLang: Efficient Execution of Structured Language Model ProgramsLianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Chuyue Sun et al.NeurIPS 2024 · 1,586 citations
- SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language ModelsGuangxuan Xiao, Ji Lin, Mickaël Seznec, Hao Wu et al.ICML 2023 · 1,493 citations
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
Related papers
- SpecCache: Speculative KV Cache Reuse for Efficient RAG ServingZijian Wen, Tao Zhang, Shuangwu Chen, Shenghao Ye et al.ACL 2026
- PKAS: Predictive KVCache-Aware Scheduling for Faster LLM and Transformer InferencesJie Ye, Avinash Maurya, Krishna Teja Chitty-Venkata, Bogdan Nicolae et al.HPDC 2026
- Cost-Efficient Large Language Model Serving for Multi-turn Conversations with CachedAttentionBin Gao, Zhuomin He, Puru Sharma, Qingxuan Kang et al.USENIX ATC 2024 · 273 citations
- LazyAttention: Efficient Retrieval-Augmented Generation with Deferred Positional EncodingHaocheng Xia, Mihir Pamnani, Hanxi Fang, Supawit Chockchowwat et al.ICML 2026
- High Throughput and Low Latency LLM Serving via Adaptive KV CachingWenyan Chen, Chengzhi Lu, Huanle Xu, Kejiang Ye et al.EuroSys 2026
