Unlocking Data-free Low-bit Quantization with Matrix Decomposition for KV Cache Compression
Peiyu Liu, Ze-Feng Gao, Xin Zhao, Yipeng Ma, Tao Wang, Ji-Rong Wen
Abstract
Key-value (KV) caching is an important technique to accelerate the inference of large language models (LLMs), but incurs significant memory overhead. To compress the size of KV cache, existing methods often compromise precision or require extra data for calibration, limiting their practicality in LLM deployment. In this paper, we introduce DecoQuant, a novel data-free low-bit quantization technique based on tensor decomposition methods, to effectively compress KV cache. Our core idea is to adjust the outlier distribution of the original matrix by performing tensor decomposition, so that the quantization difficulties are migrated from the matrix to decomposed local tensors. Specially, we find that outliers mainly concentrate on small local tensors, while large tensors tend to have a narrower value range. Based on this finding, we propose to apply low-bit quantization to the large tensor, while maintaining high-precision representation for the small tensor. Furthermore, we utilize the proposed quantization method to compress the KV cache of LLMs to accelerate the inference and develop an efficient dequantization kernel tailored specifically for DecoQuant. Through extensive experiments, DecoQuant demonstrates remarkable efficiency gains, showcasing up to a ∼75% reduction in memory footprint while maintaining comparable generation quality.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 73fd5936-cfdc-47bc-a833-e5a5272c72e5Cited by top-tier papers4
- Over-parameterized Student Model via Tensor Decomposition Boosted Knowledge DistillationYu-Liang Zhan, Zhong-Yi Lu, Hao Sun, Ze-Feng GaoNeurIPS 2024 · 6 citations
- Accurate KV Cache Eviction via Anchor Direction Projection for Efficient LLM InferenceZijie Geng, Jie Wang, Ziqi Liu, Feng Ju et al.NeurIPS 2025 · 6 citations
- LinkAlign: Scalable Schema Linking for Real-World Large-Scale Multi-Database Text-to-SQLYihan Wang, Peiyu Liu, Xin YangEMNLP 2025 · 2 citations
- ForesightKV: Optimizing KV Cache Eviction for Reasoning Models by Learning Long-Term ContributionZican Dong, Peiyu Liu, Junyi Li, Zhipeng Chen et al.ICML 2026
Builds on9
- SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language ModelsGuangxuan Xiao, Ji Lin, Mickaël Seznec, Hao Wu et al.ICML 2023 · 1,493 citations
- H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language ModelsZhenyu Zhang, Ying Sheng, Tianyi Zhou, Tianlong Chen et al.NeurIPS 2023 · 1,003 citations
- DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model ServingYinmin Zhong, Shengyu Liu, Junda Chen, Jianbo Hu et al.OSDI 2024 · 646 citations
- Scissorhands: Exploiting the Persistence of Importance Hypothesis for LLM KV Cache Compression at Test TimeZichang Liu, Aditya Desai, Fangshuo Liao, Weitao Wang et al.NeurIPS 2023 · 557 citations
- Learning to Compress Prompts with Gist TokensJesse Mu, Xiang Li, Noah D. GoodmanNeurIPS 2023 · 488 citations
Related papers
- VecInfer: Efficient LLM Inference with Low-Bit KV Cache via Outlier-Suppressed Vector QuantizationDingyu Yao, Chenxu Yang, Zhengyang Tong, Zheng Lin et al.ACL 2026 · 4 citations
- XQuant: Achieving Ultra-Low Bit KV Cache Quantization with Cross-Layer CompressionHaoqi Yang, Yao Yao, Zuchao Li, Baoyuan Qi et al.EMNLP 2025
- NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV CacheDonghyun Son, Euntae Choi, Sungjoo YooNeurIPS 2025 · 8 citations
- Accurate KV Cache Quantization with Outlier Tokens TracingYi Su, Yuechi Zhou, Quantong Qiu, Juntao Li et al.ACL 2025
- KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV CacheZirui Liu, Jiayi Yuan, Hongye Jin, Shaochen (Henry) Zhong et al.ICML 2024 · 436 citations
