EPIC: Efficient Position-Independent Caching for Serving Large Language Models
Junhao Hu, Wenrui Huang, Weidong Wang, Haoyi Wang, Tiancheng Hu, Qin Zhang, Hao Feng, Xusheng Chen, Yizhou Shan, Tao Xie
Abstract
Large Language Models (LLMs) show great capabilities in a wide range of applications, but serving them efficiently becomes increasingly challenging as requests (prompts) become more complex. Context caching improves serving performance by reusing Key-Value (KV) vectors, the intermediate representations of tokens that are repeated across requests. However, existing context caching requires exact prefix matches across requests, limiting reuse cases in settings such as few-shot learning and retrieval-augmented generation, where immutable content (e.g., documents) remains unchanged across requests but is preceded by varying prefixes. Position-Independent Caching (PIC) addresses this issue by enabling modular reuse of the KV vectors regardless of prefixes. We formalize PIC and advance prior work by introducing EPIC, a serving system incorporating our new LegoLink algorithm, which mitigates the inappropriate "attention sink" effect at every document beginning, to maintain accuracy with minimal computation. Experiments show that EPIC achieves up to 8× improvements in Time-To-First-Token (TTFT) and 7× throughput gains over existing systems, with negligible or no accuracy loss.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 968dad3c-919d-4cf3-91a0-66444687ac22Cited by top-tier papers11
- DEEPSERVE: Serverless Large Language Model Serving at ScaleJunhao Hu, Jiang Xu, Zhixia Liu, Yulong He et al.USENIX ATC 2025 · 38 citations
- ProphetKV: User-Query-Driven Selective Recomputation for Efficient KV Cache Reuse in Retrieval-Augmented GenerationShihao Wang, Jiahao Chen, Yanqi Pan, Hao Huang et al.ICML 2026 · 4 citations
- C2KV: Compressed and Composable KV Cache Reuse for Efficient LLM InferenceChuheng Du, Junyi Chen, Hanlin Tang, Kan Liu et al.KDD 2026 · 3 citations
- InfoFlow KV: Information-Flow-Aware KV Recomputation for Long ContextXin Teng, Canyu Zhang, Shaoyi Zheng, Danyang Zhuo et al.ICML 2026 · 1 citation
- Bat: Efficient Generative Recommender Serving with Bipartite AttentionJie Sun, Shaohang Wang, Zimo Zhang, Zhengyu Liu et al.ASPLOS 2026 · 1 citation
Builds on11
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han et al.ICLR 2024 · 1,714 citations
- SGLang: Efficient Execution of Structured Language Model ProgramsLianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Chuyue Sun et al.NeurIPS 2024 · 1,586 citations
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
- H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language ModelsZhenyu Zhang, Ying Sheng, Tianyi Zhou, Tianlong Chen et al.NeurIPS 2023 · 1,003 citations
- Orca: A Distributed Serving System for Transformer-Based Generative ModelsGyeong-In Yu, Joo Seong Jeong, Geon-Woo Kim, Soojeong Kim et al.OSDI 2022 · 690 citations
Related papers
- LazyAttention: Efficient Retrieval-Augmented Generation with Deferred Positional EncodingHaocheng Xia, Mihir Pamnani, Hanxi Fang, Supawit Chockchowwat et al.ICML 2026
- KVLink: Accelerating Large Language Models via Efficient KV Cache ReuseJingbo Yang, Bairu Hou, Wei Wei, Yujia Bao et al.NeurIPS 2025 · 83 citations
- Online Context Caching for Distributed Large Language Models ServingBin Gao, Zhuomin He, Yizhen Yao, Zhanzhi Lou Lou et al.INFOCOM 2025 · 2 citations
- Compute or Load KV Cache? Why Not Both?Shuowei Jin, Xueshen Liu, Qingzhao Zhang, Zhuoqing MaoICML 2025
- HijackKV: New Threat in Position-Independent KV Cache ReuseYichi Zhang, Zhiqi Wang, Huan Zhang, Yuchen YangUSENIX Security 2026
