SCVQ: Sparse-Compensated Vector Quantization for Large Language Models
Zixuan Zhou, Yujun Diao, Zicheng Kong, Dehua Ma, Zhenbo Xu, Pei Pei Li, Zhaofeng He
摘要
Large language models are primarily constrained by computing and bandwidth limitations during hardware deployment. However, current vector quantization methods incur substantial inference overhead, mainly because large-scale codebook storage and frequent index lookups consume excessive memory and compute resources. To address these challenges, we present Sparse-Compensated Vector Quantization (SCVQ), a salience and sparsecompensated vector quantization framework. By integrating a salience-aware weighted Kmeans clustering scheme with symmetry constraints, SCVQ significantly reduces codebook size and indexing costs. Crucially, we design a structured sparse representation that unifies outliers, salient weights, and quantization residuals into a single sparse matrix to to maintain model performance with minimal overhead, thereby addressing the challenges of aggressive compression. Extensive experiments across multiple benchmarks validate the strong low-bit quantization performance of SCVQ over current methods. Specifically, SCVQ achieves a perplexity of 5.87 on the WikiText-2 dataset at 2-bit quantization for LLaMA-2-7B, while delivering a 1.4× end-to-end inference speedup on an NVIDIA A100 GPU compared to the baseline.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper10
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 被引用 3,037 次
- Up or Down? Adaptive Rounding for Post-Training QuantizationMarkus Nagel, Rana Ali Amjad, Mart van Baalen, Christos Louizos 等ICML 2020 · 被引用 816 次
- QuIP: 2-Bit Quantization of Large Language Models With GuaranteesJerry Chee, Yaohui Cai, Volodymyr Kuleshov, Christopher De SaNeurIPS 2023 · 被引用 503 次
- OmniQuant: Omnidirectionally Calibrated Quantization for Large Language ModelsWenqi Shao, Mengzhao Chen, Zhaoyang Zhang, Peng Xu 等ICLR 2024 · 被引用 395 次
相关 Paper
- VecInfer: Efficient LLM Inference with Low-Bit KV Cache via Outlier-Suppressed Vector QuantizationDingyu Yao, Chenxu Yang, Zhengyang Tong, Zheng Lin 等ACL 2026 · 被引用 4 次
- VPTQ: Extreme Low-bit Vector Post-Training Quantization for Large Language ModelsYifei Liu, Jicheng Wen, Yang Wang, Shengyu Ye 等EMNLP 2024 · 被引用 10 次
- LittleBit: Ultra Low-Bit Quantization via Latent FactorizationBanseok Lee, Dongkyu Kim, Youngcheon You, Youngmin KimNeurIPS 2025 · 被引用 14 次
- SliM-LLM: Salience-Driven Mixed-Precision Quantization for Large Language ModelsWei Huang, Haotong Qin, Yangdong Liu, Yawei Li 等ICML 2025
- UniSVQ: 2-bit Unified Scalar-Vector QuantizationHaoyu Wang, Haiyan Zhao, Xingyu Yu, Zhangyang Yao 等ICML 2026 · 被引用 2 次
