ShadowKV: KV Cache in Shadows for High-Throughput Long-Context LLM Inference
Hanshi Sun, Li-Wen Chang, Wenlei Bao, Size Zheng, Ningxin Zheng, Xin Liu, Harry Dong, Yuejie Chi, Beidi Chen
摘要
With the widespread deployment of long-context large language models (LLMs), there has been a growing demand for efficient support of high-throughput inference. However, as the key-value (KV) cache expands with the sequence length, the increasing memory footprint and the need to access it for each token generation both result in low throughput when serving long-context LLMs. While various dynamic sparse attention methods have been proposed to speed up inference while maintaining generation quality, they either fail to sufficiently reduce GPU memory consumption or introduce significant decoding latency by offloading the KV cache to the CPU. We present ShadowKV, a high-throughput long-context LLM inference system that stores the low-rank key cache and offloads the value cache to reduce the memory footprint for larger batch sizes and longer sequences. To minimize decoding latency, ShadowKV employs an accurate KV selection strategy that reconstructs minimal sparse KV pairs on-the-fly. By evaluating ShadowKV on a broad range of benchmarks, including RULER, LongBench, and Needle In A Haystack, and models like Llama-3.1-8B, Llama-3-8B-1M, GLM-4-9B-1M, Yi-9B-200K, Phi-3-Mini-128K, and Qwen2-7B-128K, we demonstrate that it can support up to 6× larger batch sizes and boost throughput by up to 3.04× on an A100 GPU without sacrificing accuracy, even surpassing the performance achievable with infinite batch size under the assumption of infinite GPU memory.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper24
- Neural Attention SearchDifan Deng, Marius LindauerNeurIPS 2025 · 被引用 431 次
- R-KV: Redundancy-aware KV Cache Compression for Reasoning ModelsZefan Cai, Wen Xiao, Hanshi Sun, Cheng Luo 等NeurIPS 2025 · 被引用 50 次
- Multi-Head Low-Rank AttentionSongtao Liu, Hongwu Peng, Zhiwei Zhang, Zhengyu Chen 等ICLR 2026 · 被引用 18 次
- FreeKV: Boosting KV Cache Retrieval for Efficient LLM InferenceGuangda Liu, Chengwei Li, Zhenyu Ning, Jing Lin 等ICLR 2026 · 被引用 17 次
- ThunderAgent: A Fast, Simple, and Program-Aware Agentic Inference SystemHao Kang, Ziyang Li, Xinyu Yang, Weili Xu 等ICML 2026 · 被引用 14 次
它引用的顶会 Paper26
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra 等NeurIPS 2022 · 被引用 5,493 次
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 被引用 2,600 次
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han 等ICLR 2024 · 被引用 1,714 次
- SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language ModelsGuangxuan Xiao, Ji Lin, Mickaël Seznec, Hao Wu 等ICML 2023 · 被引用 1,493 次
- Fast Inference from Transformers via Speculative DecodingYaniv Leviathan, Matan Kalman, Yossi MatiasICML 2023 · 被引用 1,472 次
相关 Paper
- Efficient Low Rank Attention for Long-Context Inference in Large Language ModelsTenghui Li, Guoxu Zhou, Xuyang Zhao, Yuning Qiu 等NeurIPS 2025 · 被引用 4 次
- ECHO: Efficient KV Cache Offloading with Lossless Prefetching for Serving Native Sparse Attention LLMsGuangda Liu, Wenhao Chen, Chengwei Li, Zhenyu Ning 等OSDI 2026
- SCBench: A KV Cache-Centric Analysis of Long-Context MethodsYucheng Li, Huiqiang Jiang, Qianhui Wu, Xufang Luo 等ICLR 2025
- A Little Goes a Long Way: Efficient Long Context Training and Inference with Partial ContextsSuyu Ge, Xihui Lin, Yunan Zhang, Jiawei Han 等ICLR 2025
- RetroInfer: A Vector Storage Engine for Scalable Long-Context LLM InferenceYaoqi Chen, Jinkai Zhang, Baotong Lu, Qianxi Zhang 等VLDB 2026 · 被引用 1 次
