SCVQ: Sparse-Compensated Vector Quantization for Large Language Models
Zixuan Zhou, Yujun Diao, Zicheng Kong, Dehua Ma, Zhenbo Xu, Pei Pei Li, Zhaofeng He
Abstract
Large language models are primarily constrained by computing and bandwidth limitations during hardware deployment. However, current vector quantization methods incur substantial inference overhead, mainly because large-scale codebook storage and frequent index lookups consume excessive memory and compute resources. To address these challenges, we present Sparse-Compensated Vector Quantization (SCVQ), a salience and sparsecompensated vector quantization framework. By integrating a salience-aware weighted Kmeans clustering scheme with symmetry constraints, SCVQ significantly reduces codebook size and indexing costs. Crucially, we design a structured sparse representation that unifies outliers, salient weights, and quantization residuals into a single sparse matrix to to maintain model performance with minimal overhead, thereby addressing the challenges of aggressive compression. Extensive experiments across multiple benchmarks validate the strong low-bit quantization performance of SCVQ over current methods. Specifically, SCVQ achieves a perplexity of 5.87 on the WikiText-2 dataset at 2-bit quantization for LLaMA-2-7B, while delivering a 1.4× end-to-end inference speedup on an NVIDIA A100 GPU compared to the baseline.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b3f20b9b-1fc9-40cd-b25f-11f72c3532baBuilds on10
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 3,037 citations
- Up or Down? Adaptive Rounding for Post-Training QuantizationMarkus Nagel, Rana Ali Amjad, Mart van Baalen, Christos Louizos et al.ICML 2020 · 816 citations
- QuIP: 2-Bit Quantization of Large Language Models With GuaranteesJerry Chee, Yaohui Cai, Volodymyr Kuleshov, Christopher De SaNeurIPS 2023 · 503 citations
- OmniQuant: Omnidirectionally Calibrated Quantization for Large Language ModelsWenqi Shao, Mengzhao Chen, Zhaoyang Zhang, Peng Xu et al.ICLR 2024 · 395 citations
Related papers
- VecInfer: Efficient LLM Inference with Low-Bit KV Cache via Outlier-Suppressed Vector QuantizationDingyu Yao, Chenxu Yang, Zhengyang Tong, Zheng Lin et al.ACL 2026 · 4 citations
- VPTQ: Extreme Low-bit Vector Post-Training Quantization for Large Language ModelsYifei Liu, Jicheng Wen, Yang Wang, Shengyu Ye et al.EMNLP 2024 · 10 citations
- LittleBit: Ultra Low-Bit Quantization via Latent FactorizationBanseok Lee, Dongkyu Kim, Youngcheon You, Youngmin KimNeurIPS 2025 · 14 citations
- SliM-LLM: Salience-Driven Mixed-Precision Quantization for Large Language ModelsWei Huang, Haotong Qin, Yangdong Liu, Yawei Li et al.ICML 2025
- UniSVQ: 2-bit Unified Scalar-Vector QuantizationHaoyu Wang, Haiyan Zhao, Xingyu Yu, Zhangyang Yao et al.ICML 2026 · 2 citations
