S-Quant: Rethinking Weight Quantization with Seed-Based Generation
Mingzi Wang, Lancheng Zou, Shuo Yin, Zhuolun He, Bei Yu
摘要
The progressive scaling of large language models (LLMs) has consistently enhanced multimodal understanding and advanced reasoning capabilities, but has substantially increased computational and hardware execution overhead. In this paper, we present S-Quant, a novel post-method that compresses only model weights. We partition each weight tensor into fixed-size blocks and assign a single seed to each block. The seed drives a hardware-friendly Linear Feedback Shift Register (LFSR) generator that dynamically produces multiple basis matrices. Each block is then reconstructed as a linear combination of these basis matrices, with block-specific coefficients, which substantially reduces the amount of stored data, increases the data-transfer efficiency between memory and compute units, and consequently speeds up memory-bound inference for large language models. Experimental results on different LLM models ranging from 7B-70B parameters show that S-Quant attains state-of-the-art performance when weights are compressed to approximately 3-bit or 4-bit. We also design a dedicated ASIC accelerator that achieves a 4× speed-up for memory-bound LLM inference. Code is available at https://github.com/wmz-max/S-Quant.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper9
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- A Simple and Effective Pruning Approach for Large Language ModelsMingjie Sun, Zhuang Liu, Anna Bair, J. Zico KolterICLR 2024 · 被引用 794 次
- OmniQuant: Omnidirectionally Calibrated Quantization for Large Language ModelsWenqi Shao, Mengzhao Chen, Zhaoyang Zhang, Peng Xu 等ICLR 2024 · 被引用 395 次
- SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight CompressionTim Dettmers, Ruslan Svirschevski, Vage Egiazarian, Denis Kuznedelev 等ICLR 2024 · 被引用 392 次
相关 Paper
- SeedLM: Compressing LLM Weights into Seeds of Pseudo-Random GeneratorsRasoul Shafipour, David Harrison, Maxwell Horton, Jeffrey Marker 等ICLR 2025
- BitMoD: Bit-serial Mixture-of-Datatype LLM AccelerationYuzong Chen, Ahmed F. AbouElhamayed, Xilai Dai, Yang Wang 等HPCA 2025 · 被引用 23 次
- NanoQuant: Efficient Sub-1-Bit Quantization of Large Language ModelsHyochan Chong, Dongkyu Kim, Changdong Kim, Minseop ChoiICML 2026
- UniSVQ: 2-bit Unified Scalar-Vector QuantizationHaoyu Wang, Haiyan Zhao, Xingyu Yu, Zhangyang Yao 等ICML 2026 · 被引用 2 次
- SEPTQ: A Simple and Effective Post-Training Quantization Paradigm for Large Language ModelsHan Liu, Haotian Gao, Xiaotong Zhang, Changya Li 等KDD 2025 · 被引用 1 次
