SpinQuant: LLM Quantization with Learned Rotations
Zechun Liu, Changsheng Zhao, Igor Fedorov, Bilge Soran, Dhruv Choudhary, Raghuraman Krishnamoorthi, Vikas Chandra, Yuandong Tian, Tijmen Blankevoort
摘要
Post-training quantization (PTQ) techniques applied to weights, activations, and the KV cache greatly reduce memory usage, latency, and power consumption of Large Language Models (LLMs), but may lead to large quantization errors when outliers are present. Rotating activation or weight matrices helps remove outliers and benefits quantization. In this work, we identify a collection of applicable rotation parameterizations that lead to identical outputs in full-precision Transformer architectures while enhancing quantization accuracy. In addition, we find that some random rotations lead to much better quantization than others, with an up to 13 points difference in downstream zero-shot reasoning performance. As a result, we propose SpinQuant, a novel approach that incorporates learned rotation matrices for optimal quantized network accuracy. With 4-bit quantization of weight, activation, and KV-cache, SpinQuant narrows the accuracy gap on zero-shot reasoning tasks with full precision to merely 2.9 points on the LLaMA-2 7B model, surpassing LLM-QAT by 19.1 points and SmoothQuant by 25.0 points. Furthermore, Spin-Quant also outperforms concurrent work QuaRot, which applies random rotations to remove outliers. In particular, for LLaMA-3 8B models that are hard to quantize, SpinQuant reduces the gap to full precision by up to 45.1% relative to QuaRot. Code is available at github.com/facebookresearch/SpinQuant.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper99
- ParetoQ: Improving Scaling Laws in Extremely Low-bit LLM QuantizationZechun Liu, Changsheng Zhao, Hanxian Huang, Sijia Chen 等NeurIPS 2025 · 被引用 50 次
- BitsFusion: 1.99 bits Weight Quantization of Diffusion ModelYang Sui, Yanyu Li, Anil Kag, Yerlan Idelbayev 等NeurIPS 2024 · 被引用 48 次
- Bridging the Gap Between Promise and Performance for Microscaling FP4 QuantizationVage Egiazarian, Roberto L. Castro, Denis Kuznedelev, Andrei Panferov 等ICLR 2026 · 被引用 38 次
- Qronos: Correcting the Past by Shaping the Future... in Post-Training QuantizationShihao Zhang, Haoyu Zhang, Ian Colbert, Rayan SaabICLR 2026 · 被引用 27 次
- Understanding and Minimising Outlier Features in Transformer TrainingBobby He, Lorenzo Noci, Daniele Paliotta, Imanol Schlag 等NeurIPS 2024 · 被引用 27 次
它引用的顶会 Paper19
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 被引用 3,037 次
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao 等AAAI 2020 · 被引用 2,916 次
- SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language ModelsGuangxuan Xiao, Ji Lin, Mickaël Seznec, Hao Wu 等ICML 2023 · 被引用 1,493 次
- Up or Down? Adaptive Rounding for Post-Training QuantizationMarkus Nagel, Rana Ali Amjad, Mart van Baalen, Christos Louizos 等ICML 2020 · 被引用 816 次
- QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMsSaleh Ashkboos, Amirkeivan Mohtashami, Maximilian L. Croci, Bo Li 等NeurIPS 2024 · 被引用 723 次
相关 Paper
- ReSpinQuant: Efficient Layer-Wise LLM Quantization via Subspace Residual Rotation ApproximationSuyoung Kim, Sunghyun Wee, Hyeonjin Kim, Kyomin Hwang 等ICML 2026 · 被引用 1 次
- ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank ResidualsUtkarsh Saxena, Sayeh Sharify, Kaushik Roy, Xin WangICML 2025
- FlatQuant: Flatness Matters for LLM QuantizationYuxuan Sun, Ruikang Liu, Haoli Bai, Han Bao 等ICML 2025
- NestQuant: nested lattice quantization for matrix products and LLMsSemyon Savkin, Eitan Porat, Or Ordentlich, Yury PolyanskiyICML 2025
- ParoQuant: Pairwise Rotation Quantization for Efficient Reasoning LLM InferenceYesheng Liang, Haisheng Chen, Song Han, Zhijian LiuICLR 2026 · 被引用 19 次
