vAttention: Dynamic Memory Management for Serving LLMs without PagedAttention
Ramya Prabhu, Ajay Nayak, Jayashree Mohan, Ramachandran Ramjee, Ashish Panwar
摘要
PagedAttention is a popular approach for dynamic memory allocation in LLM serving systems. It enables on-demand allocation of GPU memory to mitigate KV cache fragmentation -a phenomenon that crippled the batch size (and consequently throughput) in prior systems. However, in trying to allocate physical memory at runtime, PagedAttention ends up changing the virtual memory layout of the KV cache from contiguous to non-contiguous. Such a design leads to non-trivial programming and performance overheads.
We present vAttention -an approach that mitigates fragmentation in physical memory while retaining the virtual memory contiguity of the KV cache. We achieve this by decoupling the allocation of virtual and physical memory using CUDA virtual memory management APIs. We also introduce various LLM-specific optimizations to address the limitations of CUDA virtual memory support. Overall, vAttention is a simpler, portable, and performant alternative to PagedAttention: it supports various attention kernels outof-the-box and improves LLM serving throughput by up to 1.23× compared to the use of PagedAttention-based kernels of FlashAttention-2 and FlashInfer.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper24
- Prism: Cost-Efficient Multi-LLM Serving via GPU Memory BallooningShan Yu, Yifan Qiao, Mingyuan Ma, Yangmin Li 等OSDI 2026 · 被引用 33 次
- POD-Attention: Unlocking Full Prefill-Decode Overlap for Faster LLM InferenceAditya K. Kamath, Ramya Prabhu, Jayashree Mohan, Simon Peter 等ASPLOS 2025 · 被引用 24 次
- Efficient LLM Serving for Agentic Workflows: A Data Systems PerspectiveNoppanat Wadlom, Junyi Shen, Yao LuSIGMOD 2026 · 被引用 14 次
- DualMap: Enabling Both Cache Affinity and Load Balancing for Distributed LLM ServingYing Yuan, Pengfei Zuo, Bo Wang, Zhangyu Chen 等ICLR 2026 · 被引用 10 次
- Early Silicon of Raptor: The First 3D-DRAM Accelerator for Generative InferencePrashant J. Nair, Ramyad Hadidi, Subramani Ganesh, Sangamesh Kodge 等ISCA 2026 · 被引用 4 次
它引用的顶会 Paper23
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra 等NeurIPS 2022 · 被引用 5,493 次
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 被引用 2,600 次
- SGLang: Efficient Execution of Structured Language Model ProgramsLianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Chuyue Sun 等NeurIPS 2024 · 被引用 1,586 次
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
- H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language ModelsZhenyu Zhang, Ying Sheng, Tianyi Zhou, Tianlong Chen 等NeurIPS 2023 · 被引用 1,003 次
相关 Paper
- ConServe: Contiguity-Preserving Memory Management for Multi-Turn LLM ServingBingyao LiISCA 2026
- Jenga: Effective Memory Management for Serving LLM with HeterogeneityChen Zhang, Kuntai Du, Shu Liu, Woosuk Kwon 等SOSP 2025
- PAT: Accelerating LLM Decoding via Prefix-Aware Attention with Resource Efficient Multi-Tile KernelJinjun Yi, Zhixin Zhao, Yitao Hu, Ke Yan 等ASPLOS 2026 · 被引用 1 次
- Cost-Efficient Large Language Model Serving for Multi-turn Conversations with CachedAttentionBin Gao, Zhuomin He, Puru Sharma, Qingxuan Kang 等USENIX ATC 2024 · 被引用 273 次
- Accelerating LLM Inference Throughput via Asynchronous KV Cache PrefetchingYanhao Dong, Yubo Miao, Weinan Li, Xiao Zheng 等AAAI 2026 · 被引用 4 次
