WaterSIC: information-theoretically (near) optimal linear layer quantization
Egor Lifar, Semyon Savkin, Or Ordentlich, Yury Polyanskiy
摘要
This paper considers the problem of converting a given dense linear layer into a low-precision version. The tradeoff between minimizing description length and discrepancy introduced at the output of the layer is analyzed information theoretically (IT). It is shown that the popular GPTQ algorithm may have an arbitrarily large gap to IT limit. To alleviate this problem a novel algorithm, termed ''WaterSIC'', is proposed and is shown to be within a rate gap of 0.255 bit to IT limit, uniformly over all possible covariance matrices of input activations. WaterSIC's key innovation is allocating different quantization rates to different columns (in-features) of the weight matrix, mimicking the classical IT solution known as ''waterfilling''. Applying WaterSIC to real LLMs establishes new state-of-the-art for rates in the range of 1...4 bits per entry.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper20
- QuIP: 2-Bit Quantization of Large Language Models With GuaranteesJerry Chee, Yaohui Cai, Volodymyr Kuleshov, Christopher De SaNeurIPS 2023 · 被引用 503 次
- OmniQuant: Omnidirectionally Calibrated Quantization for Large Language ModelsWenqi Shao, Mengzhao Chen, Zhaoyang Zhang, Peng Xu 等ICLR 2024 · 被引用 395 次
- QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice CodebooksAlbert Tseng, Jerry Chee, Qingyao Sun, Volodymyr Kuleshov 等ICML 2024 · 被引用 295 次
- Extreme Compression of Large Language Models via Additive QuantizationVage Egiazarian, Andrei Panferov, Denis Kuznedelev, Elias Frantar 等ICML 2024 · 被引用 187 次
- QTIP: Quantization with Trellises and Incoherence ProcessingAlbert Tseng, Qingyao Sun, David Hou, Christopher De SaNeurIPS 2024 · 被引用 101 次
相关 Paper
- NestQuant: nested lattice quantization for matrix products and LLMsSemyon Savkin, Eitan Porat, Or Ordentlich, Yury PolyanskiyICML 2025
- ASER: Activation Smoothing and Error Reconstruction for Large Language Model QuantizationWeibo Zhao, Yubin Shi, Xinyu Lyu, Wanchen Sui 等AAAI 2025 · 被引用 7 次
- The Lattice Geometry of Neural Network Quantization: A Short Equivalence Proof of GPTQ and Babai's AlgorithmJohann BirnickICLR 2026 · 被引用 12 次
- Rethinking Residual Errors in Compensation-based LLM QuantizationShuaiting Li, Juncan Deng, Kedong Xu, Rongtao Deng 等ICLR 2026 · 被引用 4 次
- GPTAQ: Efficient Finetuning-Free Quantization for Asymmetric CalibrationYuhang Li, Ruokai Yin, Donghyun Lee, Shiting Xiao 等ICML 2025
