Accelerating LLM Serving for Multi-turn Dialogues with Efficient Resource Management
Jinwoo Jeong, Jeongseob Ahn
Abstract
Although there have been significant efforts to make LLM serving efficient, we observe two limitations of current state-of-the-art serving frameworks in handling multi-turn dialogues between users and assistants, particularly in chat scenarios. First, existing LLM frameworks incur substantial computational overhead in recomputing attention keys and values (KVs) for understanding context across multiple turns of user queries. Second, as the prompt length of user queries is amplified due to multi-turns, a first-come-first-served (FCFS) scheduling policy often causes head-of-line blocking issues, leading to underutilization of GPU resources.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers10
- Strata: Hierarchical Context Caching for Long Context Language Model ServingZhiqiang Xie, Ziyi Xu, Mark Zhao, Yuwei An et al.OSDI 2026 · 40 citations
- The Cost of Dynamic Reasoning: Demystifying AI Agents and Test-Time Scaling from an AI Infrastructure PerspectiveJiin Kim, Byeongjun Shin, Jinha Chung, Minsoo RhuHPCA 2026 · 7 citations
- KVDrive: A Holistic Multi-Tier KV Cache Management System for Long-Context LLM InferenceJian Lin, Jiazhi Mi, Zicong Hong, Haodong Wang et al.SIGMOD 2026 · 5 citations
- AugServe: Adaptive Request Scheduling for Augmented Large Language Model Inference ServingYing Wang, Zhen Jin, Zhenqian Chen, Jiexiong Xu et al.ICML 2026 · 4 citations
- Bidaw: Enhancing Key-Value Caching for Interactive LLM Serving via Bidirectional Computation-Storage AwarenessShipeng Hu, Guangyan Zhang, Yuqi Zhou, Yaya Wei et al.FAST 2026 · 3 citations
Related papers
- Cost-Efficient Large Language Model Serving for Multi-turn Conversations with CachedAttentionBin Gao, Zhuomin He, Puru Sharma, Qingxuan Kang et al.USENIX ATC 2024 · 273 citations
- Stateful Large Language Model Serving with PensieveLingfan Yu, Jinkun Lin, Jinyang LiEuroSys 2025 · 23 citations
- Aqua: Network-Accelerated Memory Offloading for LLMs in Scale-Up GPU DomainsAbhishek Vijaya Kumar, Gianni Antichi, Rachee SinghASPLOS 2025 · 4 citations
- ConServe: Contiguity-Preserving Memory Management for Multi-Turn LLM ServingBingyao LiISCA 2026
- FastServe: Iteration-Level Preemptive Scheduling for Large Language Model InferenceBingyang Wu, Yinmin Zhong, Zili Zhang, Shengyu Liu et al.NSDI 2026 · 12 citations
