ConServe: Contiguity-Preserving Memory Management for Multi-Turn LLM Serving
Bingyao Li
Abstract
Interacting with humans through multi-turn conversations is a fundamental capability of large language models (LLMs). Multi-turn serving stresses GPU memory systems as KV caches grow, persist, and are heavily reused across turns. PagedAttention, the prevailing dynamic allocator in modern systems, mitigates fragmentation and increases batch size by managing the KV cache in blocks, but in doing so converts an originally contiguous virtual address space into a non-contiguous layout, introducing non-trivial programming complexity and additional address translation overheads from scattered virtual pages and block table lookups. We reveal that translation locality is a critical bottleneck under such paged KV designs, and we present ConServe, a contiguity-aware virtual memory allocator that retains per-conversation contiguity in the virtual address space while mitigating fragmentation in physical memory. Con-Serve reserves a single contiguous virtual address slice per conversation, maps physical pages on demand with CUDA VMM, and resizes virtual address slices via lazy, copy-free remapping at negligible overhead. Across diverse models and workloads, ConServe achieves up to 74.4%,43.4%, and 19.1% lower TTFT and 35.1%, 15.6%, and 12.1% higher end-to-end throughput than vLLM, vAttention-Turn, and vAttention-Conv, respectively.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Related papers
- vAttention: Dynamic Memory Management for Serving LLMs without PagedAttentionRamya Prabhu, Ajay Nayak, Jayashree Mohan, Ramachandran Ramjee et al.ASPLOS 2025 · 38 citations
- Stateful Large Language Model Serving with PensieveLingfan Yu, Jinkun Lin, Jinyang LiEuroSys 2025 · 23 citations
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
- ConServe: Fine-Grained GPU Harvesting for LLM Online and Offline Co-ServingYifan Qiao, Shan Yu, Shu Anzai, Haoran Ma et al.ICML 2026 · 7 citations
- Cost-Efficient Large Language Model Serving for Multi-turn Conversations with CachedAttentionBin Gao, Zhuomin He, Puru Sharma, Qingxuan Kang et al.USENIX ATC 2024 · 273 citations
