PQCache: Product Quantization-based KVCache for Long Context LLM Inference
Hailin Zhang, Xiaodong Ji, Yilin Chen, Fangcheng Fu, Xupeng Miao, Xiaonan Nie, Weipeng Chen, Bin Cui
Abstract
As the field of Large Language Models (LLMs) continues to evolve, the context length in inference is steadily growing. Key-Value Cache (KVCache), the intermediate representations of tokens within LLM inference, has now become the primary memory bottleneck due to limited GPU memory. Current methods selectively determine suitable keys and values for self-attention computation in LLMs to address the issue. However, they either fall short in maintaining model quality or result in high serving latency. Drawing inspiration from advanced embedding retrieval techniques prevalent in the data management community, we consider the storage and retrieval of KVCache as a typical embedding retrieval problem. We propose PQCache , which employs Product Quantization (PQ) to manage KVCache, maintaining model quality while ensuring low serving latency. During the prefilling phase, we apply PQ to tokens' keys for each LLM layer and head. During the autoregressive decoding phase, we use PQ codes and centroids to approximately identify important preceding tokens, then fetch the corresponding key-value pairs for self-attention computation. Through meticulous design of overlapping and caching, we minimize any additional computation and communication overhead during both phases. Extensive experiments demonstrate that PQCache achieves both effectiveness and efficiency, with 4.60% score improvement over existing methods on InfiniteBench and low system latency in both prefilling and decoding.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d5ff44eb-e4bb-42ee-804f-f4eb64117672Cited by top-tier papers39
- Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM InferenceYuan Feng, Junlin Lv, Yukun Cao, Xike Xie et al.NeurIPS 2025 · 256 citations
- RetrievalAttention: Accelerating Long-Context LLM Inference via Vector RetrievalDi Liu, Meng Chen, Baotong Lu, Huiqiang Jiang et al.NeurIPS 2025 · 148 citations
- Twilight: Adaptive Attention Sparsity with Hierarchical Top- PruningChaofan Lin, Jiaming Tang, Shuo Yang, Hanshuo Wang et al.NeurIPS 2025 · 53 citations
- Don’t Pass@k: A Bayesian Framework for Large Language Model EvaluationMohsen Hariri, Amirhossein Samandar, Michael Hinczewski, Vipin ChaudharyICLR 2026 · 18 citations
- Sparse Attention Adaptation for Long ReasoningYizhao Gao, Shuming Guo, Shijie Cao, Yuqing Xia et al.ICLR 2026 · 17 citations
Builds on49
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 2,600 citations
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han et al.ICLR 2024 · 1,714 citations
Related papers
- MILLION: MasterIng Long-Context LLM Inference Via Outlier-Immunized KV Product QuaNtizationZongwu Wang, Peng Xu, Fangxin Liu, Yiwei Hu et al.DAC 2025 · 6 citations
- VecInfer: Efficient LLM Inference with Low-Bit KV Cache via Outlier-Suppressed Vector QuantizationDingyu Yao, Chenxu Yang, Zhengyang Tong, Zheng Lin et al.ACL 2026 · 4 citations
- IceCache: Memory-Efficient KV-cache Management for Long-Sequence LLMsYuzhen Mao, Qitong Wang, Martin Ester, Ke LiICLR 2026 · 6 citations
- Efficient Cooperation-Aware Key and Value Management for LLM InferenceQiheng Sun, Hongwei Zhang, Junxu Liu, Haocheng Xia et al.VLDB 2026
- JanusQuant: Accurate and Efficient 2-bit KV Cache Quantization for Long-Context InferenceChengyu Sun, Yaqi Xia, Hulin Wang, Donglin Yang et al.PPoPP 2026 · 1 citation
