The Lattice Geometry of Neural Network Quantization: A Short Equivalence Proof of GPTQ and Babai's Algorithm
Johann Birnick
摘要
We explain how data-driven quantization of a linear unit in a neural network corresponds to solving the closest vector problem for a certain lattice generated by input data. We prove that the GPTQ algorithm (Frantar et al., 2023) is equivalent to Babai's well-known nearest-plane algorithm (Babai, 1986) . We furthermore provide geometric intuition for both algorithms. Lastly, we note the consequences of these results, in particular hinting at the possibility of using lattice basis reduction for improved quantization. QUANTIZATION AND LATTICES Computations in neural networks are usually carried out in 32-bit or 16-bit floating point arithmetic. In particular, the parameters (weights) of the network are stored in this comparatively high precision. Quantization is the art of reducing precision, in favor of less memory consumption and faster computation, while keeping the accuracy as high as possible. In this paper, we are interested only in post-training quantization of the weights: We are handed a trained neural network, and our goal is to approximate (some of) the parameters of the network with a coarse numerical alphabet, while keeping the accuracy high. Commonly, this effort is focused on the linear parts of the network. That is, we are given a linear map R n → R m , represented by a weight matrix W ∈ R m×n , and we seek to find another m × n matrix V , whose entries have lower numerical precision and which "approximates W well". Concretely, this means:
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Qronos: Correcting the Past by Shaping the Future... in Post-Training QuantizationShihao Zhang, Haoyu Zhang, Ian Colbert, Rayan SaabICLR 2026 · 被引用 27 次
- The Geometry of LLM Quantization: GPTQ as Babai's Nearest Plane AlgorithmJiale Chen, Yalda Shabanzadeh, Elvir Crnčević, Torsten Hoefler 等ICLR 2026 · 被引用 27 次
- WaterSIC: information-theoretically (near) optimal linear layer quantizationEgor Lifar, Semyon Savkin, Or Ordentlich, Yury PolyanskiyICML 2026 · 被引用 5 次
它引用的顶会 Paper4
- Optimal Brain Compression: A Framework for Accurate Post-Training Quantization and PruningElias Frantar, Dan AlistarhNeurIPS 2022 · 被引用 440 次
- Qronos: Correcting the Past by Shaping the Future... in Post-Training QuantizationShihao Zhang, Haoyu Zhang, Ian Colbert, Rayan SaabICLR 2026 · 被引用 27 次
- The Geometry of LLM Quantization: GPTQ as Babai's Nearest Plane AlgorithmJiale Chen, Yalda Shabanzadeh, Elvir Crnčević, Torsten Hoefler 等ICLR 2026 · 被引用 27 次
- OPTQ: Accurate Quantization for Generative Pre-trained TransformersElias Frantar, Saleh Ashkboos, Torsten Hoefler, Dan AlistarhICLR 2023
相关 Paper
- NestQuant: nested lattice quantization for matrix products and LLMsSemyon Savkin, Eitan Porat, Or Ordentlich, Yury PolyanskiyICML 2025
- Learning Grouped Lattice Vector Quantizers for Low-Bit LLM CompressionXi Zhang, Xiaolin Wu, Jiamang Wang, Weisi LinNeurIPS 2025 · 被引用 4 次
- Up or Down? Adaptive Rounding for Post-Training QuantizationMarkus Nagel, Rana Ali Amjad, Mart van Baalen, Christos Louizos 等ICML 2020 · 被引用 816 次
- Demystifying and Generalizing BinaryConnectTim Dockhorn, Yaoliang Yu, Eyyüb Sari, Mahdi Zolnouri 等NeurIPS 2021 · 被引用 14 次
- OMPQ: Orthogonal Mixed Precision QuantizationYuexiao Ma, Taisong Jin, Xiawu Zheng, Yan Wang 等AAAI 2023 · 被引用 56 次
