Lune

ACL2026Top-tier venue

SCVQ: Sparse-Compensated Vector Quantization for Large Language Models

Zixuan Zhou, Yujun Diao, Zicheng Kong, Dehua Ma, Zhenbo Xu, Pei Pei Li, Zhaofeng He

2026Year

Abstract

Large language models are primarily constrained by computing and bandwidth limitations during hardware deployment. However, current vector quantization methods incur substantial inference overhead, mainly because large-scale codebook storage and frequent index lookups consume excessive memory and compute resources. To address these challenges, we present Sparse-Compensated Vector Quantization (SCVQ), a salience and sparsecompensated vector quantization framework. By integrating a salience-aware weighted Kmeans clustering scheme with symmetry constraints, SCVQ significantly reduces codebook size and indexing costs. Crucially, we design a structured sparse representation that unifies outliers, salient weights, and quantization residuals into a single sparse matrix to to maintain model performance with minimal overhead, thereby addressing the challenges of aggressive compression. Extensive experiments across multiple benchmarks validate the strong low-bit quantization performance of SCVQ over current methods. Specifically, SCVQ achieves a perplexity of 5.87 on the WikiText-2 dataset at 2-bit quantization for LLaMA-2-7B, while delivering a 1.4× end-to-end inference speedup on an NVIDIA A100 GPU compared to the baseline.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext b3f20b9b-1fc9-40cd-b25f-11f72c3532ba

Builds on10

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines