RelayCaching: Accelerating LLM Collaboration via Decoding KV Cache Reuse
Yingsheng Geng, Yuchong Gao, Weihong Wu, Guyue Liu, Jiang liu
Abstract
The increasing complexity of AI tasks has shifted the paradigm from monolithic models toward multi-agent large language model (LLM) systems. However, these collaborative architectures introduce a critical bottleneck: redundant prefill computation for shared content generated by previous agents, which significantly increases KV cache memory usage and time-to-first-token (TTFT). While various KV cache methods have been proposed to mitigate prefill redundancy, they either fail to maintain accuracy on agent-generated outputs or exhibit low reuse rates due to rigid constraints. We present RelayCaching, a training-free inference method that directly reuses decoding phase KV caches from previous agents in subsequent prefill phases. Our key insight is that KV caches for identical content are highly consistent across phases, while prefix-induced deviations are sparse and localized within a limited range of layers and token positions. By selectively recomputing KV caches at these positions, RelayCaching preserves model accuracy with minimal overhead, yielding a superior accuracy–efficiency trade-off over existing methods. Experiments on diverse collaborative LLM tasks spanning mathematical reasoning, general knowledge, and code generation demonstrate that RelayCaching achieves over % KV cache reuse, reduces TTFT by up to compared to the standard pipeline, all with negligible accuracy degradation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b5c470d4-57c0-447f-a0b7-9caf4e70ba10Builds on11
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- CAMEL: Communicative Agents for "Mind" Exploration of Large Language Model SocietyGuohao Li, Hasan Hammoud, Hani Itani, Dmitrii Khizbullin et al.NeurIPS 2023 · 1,975 citations
- SGLang: Efficient Execution of Structured Language Model ProgramsLianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Chuyue Sun et al.NeurIPS 2024 · 1,586 citations
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
- Many Heads Are Better Than One: Improved Scientific Idea Generation by A LLM-Based Multi-Agent SystemHaoyang Su, Renqi Chen, Shixiang Tang, Zhenfei Yin et al.ACL 2025 · 49 citations
Related papers
- KVCOMM: Online Cross-context KV-cache Communication for Efficient LLM-based Multi-agent SystemsHancheng Ye, Zhengqi Gao, Mingyuan Ma, Qinsi Wang et al.NeurIPS 2025 · 42 citations
- ICaRus: Identical Cache Reuse for Efficient Multi-Model InferenceSunghyeon Woo, Jaeeun Kil, Hoseung Kim, Minsub Kim et al.ICLR 2026 · 7 citations
- DroidSpeak: KV Cache Sharing Across Fine-tuned Model VariantsYuhan Liu, Yuyang Huang, Jiayi Yao, Shaoting Feng et al.NSDI 2026 · 14 citations
- LazyAttention: Efficient Retrieval-Augmented Generation with Deferred Positional EncodingHaocheng Xia, Mihir Pamnani, Hanxi Fang, Supawit Chockchowwat et al.ICML 2026
- Attention Is All You Need for KV Cache in Diffusion LLMsQuan Nguyen-Tri, Mukul Ranjan, Zhiqiang ShenICLR 2026 · 36 citations
