Lune

SIGCOMM2026Top-tier venue

DualPath: Accelerating Agentic LLM Inference by Harvesting Disaggregated KV-Cache Storage I/O

Yongtong Wu, Shaoyuan Chen, Rilin Huang, Yixuan Tan, Yinmin Zhong, Mingxing Zhang, Xin Jin, Panpan Huang

2026Year

Abstract

The performance of multi-turn, agentic LLM inference is increasingly dominated by KV-Cache storage I/O rather than computation. In prevalent disaggregated architectures, loading the massive KV-Cache from external storage creates a fundamental imbalance: storage NICs on prefill engines become bandwidth-saturated, while those on decoding engines remain idle. This asymmetry severely constrains overall system throughput.

Ask about this paper

Ask your agent about it.

Lune has read the top-tier papers around this one, so every answer names the papers it rests on.

Questions to start from

Your agent calls

Lunesearch_papers

Ask in Lune

Free to start. No credit card required.

lune papers get 386e820d-514e-406a-9df9-51bda7b484f8

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines