ParoQuant: Pairwise Rotation Quantization for Efficient Reasoning LLM Inference
Yesheng Liang, Haisheng Chen, Song Han, Zhijian Liu
摘要
Post-training quantization (PTQ) compresses the weights and activations of large language models (LLMs) into low-precision representations to reduce memory footprint and accelerate inference. However, the presence of outliers in weights and activations often leads to large quantization errors and severe accuracy degradation, especially in recent reasoning LLMs where errors accumulate across long chains of thought. Existing PTQ methods either fail to sufficiently suppress outliers or introduce significant overhead during inference. In this paper, we propose Pairwise Rotation Quantization (ParoQuant), a PTQ method that combines hardware-efficient and optimizable independent Givens rotations with channel-wise scaling to even out the magnitudes across channels and narrow the dynamic range within each quantization group, effectively addressing the outlier issue. We further co-design the inference kernel to fully exploit GPU parallelism and keep the rotations and scaling lightweight at runtime. Under weight-only quantization, ParoQuant achieves an average 2.4% accuracy improvement over AWQ on reasoning tasks, with less than 10% overhead. ParoQuant also matches the accuracy of state-of-the-art weight-activation quantization methods. This paves the way for more efficient and accurate deployment of reasoning LLMs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- ReSpinQuant: Efficient Layer-Wise LLM Quantization via Subspace Residual Rotation ApproximationSuyoung Kim, Sunghyun Wee, Hyeonjin Kim, Kyomin Hwang 等ICML 2026 · 被引用 1 次
- Proteus: Lookup-Free Trellis-Coded Quantization by Lattice-Breaking Compute Codes for 2-Bit LLMsZhengwu Yang, Xunchao Li, Ke Cheng, Kunlong Liu 等ICML 2026
它引用的顶会 Paper18
- SGLang: Efficient Execution of Structured Language Model ProgramsLianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Chuyue Sun 等NeurIPS 2024 · 被引用 1,586 次
- SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language ModelsGuangxuan Xiao, Ji Lin, Mickaël Seznec, Hao Wu 等ICML 2023 · 被引用 1,493 次
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
- QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMsSaleh Ashkboos, Amirkeivan Mohtashami, Maximilian L. Croci, Bo Li 等NeurIPS 2024 · 被引用 723 次
- QuIP: 2-Bit Quantization of Large Language Models With GuaranteesJerry Chee, Yaohui Cai, Volodymyr Kuleshov, Christopher De SaNeurIPS 2023 · 被引用 503 次
相关 Paper
- ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank ResidualsUtkarsh Saxena, Sayeh Sharify, Kaushik Roy, Xin WangICML 2025
- RUQuant: Towards Refining Uniform Quantization for Large Language ModelsHan Liu, Haotian Gao, Changya Li, Feng Zhang 等KDD 2026
- Pushing the Limits of Block Rotations in Post-Training QuantizationSai Sanjeet, Ian Colbert, Pablo Monteagudo-Lago, Giuseppe Franco 等ICML 2026
- SpinQuant: LLM Quantization with Learned RotationsZechun Liu, Changsheng Zhao, Igor Fedorov, Bilge Soran 等ICLR 2025
- TWLA: Achieving Ternary Weights and Low-Bit Activations for LLMs via Post-Training QuantizationZhixiong Zhao, Zukang Xu, Zhixuan Chen, Xing Hu 等ICML 2026
