DualPath: Accelerating Agentic LLM Inference by Harvesting Disaggregated KV-Cache Storage I/O
Yongtong Wu, Shaoyuan Chen, Rilin Huang, Yixuan Tan, Yinmin Zhong, Mingxing Zhang, Xin Jin, Panpan Huang
2026Year
Abstract
The performance of multi-turn, agentic LLM inference is increasingly dominated by KV-Cache storage I/O rather than computation. In prevalent disaggregated architectures, loading the massive KV-Cache from external storage creates a fundamental imbalance: storage NICs on prefill engines become bandwidth-saturated, while those on decoding engines remain idle. This asymmetry severely constrains overall system throughput.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 386e820d-514e-406a-9df9-51bda7b484f8Related papers
- Not All Prefills Are Equal: PPD Disaggregation for Multi-turn LLM ServingZongze Li, Jingyu Liu, Zach Xu, Yineng Zhang et al.ICML 2026 · 4 citations
- ICaRus: Identical Cache Reuse for Efficient Multi-Model InferenceSunghyeon Woo, Jaeeun Kil, Hoseung Kim, Minsub Kim et al.ICLR 2026 · 7 citations
- RelayCaching: Accelerating LLM Collaboration via Decoding KV Cache ReuseYingsheng Geng, Yuchong Gao, Weihong Wu, Guyue Liu et al.ICML 2026
- DroidSpeak: KV Cache Sharing Across Fine-tuned Model VariantsYuhan Liu, Yuyang Huang, Jiayi Yao, Shaoting Feng et al.NSDI 2026 · 14 citations
- PKAS: Predictive KVCache-Aware Scheduling for Faster LLM and Transformer InferencesJie Ye, Avinash Maurya, Krishna Teja Chitty-Venkata, Bogdan Nicolae et al.HPDC 2026
