FPTQuant: Function-Preserving Transforms for LLM Quantization
Boris van Breugel, Yelysei Bondarenko, Paul Whatmough, Markus Nagel
摘要
Large language models (LLMs) require substantial compute, and thus energy, at inference time. While quantizing weights and activations is effective at improving efficiency, naive quantization of LLMs can significantly degrade performance due to large magnitude outliers. This paper describes FPTQuant, which introduces three novel, lightweight, and expressive function-preserving transforms (FPTs) to facilitate quantization of transformers: (1) a mergeable pre-RoPE transform for queries and keys, (2) a mergeable transform for values, and (3) a cheap, dynamic pertoken scaling transform. By leveraging the equivariances and independencies inherent to canonical transformer operation, we designed these FPTs to maintain the model's function while shaping the intermediate activation distributions to be more quantization friendly. FPTQuant requires no custom kernels and adds virtually no overhead during inference. The FPTs are trained both locally to reduce outliers, and end-to-end such that the outputs of the quantized and full-precision models match. FPTQuant enables static INT4 quantization with minimal overhead and shows SOTA speed-up of up to 3.9× over FP. Empirically, FPTQuant has an excellent accuracy-speed trade-off-it is performing on par or exceeding most prior work and only shows slightly lower accuracy compared to a method that is up to 29% slower.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Block Rotation is All You Need for MXFP4 QuantizationYuantian Shao, Peisong Wang, Yuanteng Chen, Chang Xu 等ICML 2026 · 被引用 16 次
- DeepProve: Verifiable End-to-End Large Language Model InferenceNicolas Gailly, Ismael Hishon-Rezaizadeh, Tianyi Liu, Nicholas Mainardi 等CCS 2026
- STaMP: Sequence Transformation and Mixed Precision for Low-Precision Activation QuantizationMarco Federici, Riccardo Del Chiaro, Boris van Breugel, Paul N. Whatmough 等ICLR 2026
它引用的顶会 Paper24
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 被引用 3,037 次
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao 等AAAI 2020 · 被引用 2,916 次
- SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language ModelsGuangxuan Xiao, Ji Lin, Mickaël Seznec, Hao Wu 等ICML 2023 · 被引用 1,493 次
- Learned Step Size quantizationSteven K. Esser, Jeffrey L. McKinstry, Deepika Bablani, Rathinakumar Appuswamy 等ICLR 2020 · 被引用 1,037 次
- Up or Down? Adaptive Rounding for Post-Training QuantizationMarkus Nagel, Rana Ali Amjad, Mart van Baalen, Christos Louizos 等ICML 2020 · 被引用 816 次
相关 Paper
- Understanding Int4 Quantization for Language Models: Latency Speedup, Composability, and Failure CasesXiaoxia Wu, Cheng Li, Reza Yazdani Aminabadi, Zhewei Yao 等ICML 2023 · 被引用 41 次
- LLM-FP4: 4-Bit Floating-Point Quantized TransformersShih-Yang Liu, Zechun Liu, Xijie Huang, Pingcheng Dong 等EMNLP 2023 · 被引用 34 次
- HALO: Hadamard-Assisted Lower-Precision Optimization for LLMsSaleh Ashkboos, Mahdi Nikdan, Rush Tabesh, Roberto L. Castro 等NeurIPS 2025 · 被引用 16 次
- Bridging the Gap Between Promise and Performance for Microscaling FP4 QuantizationVage Egiazarian, Roberto L. Castro, Denis Kuznedelev, Andrei Panferov 等ICLR 2026 · 被引用 38 次
- ParoQuant: Pairwise Rotation Quantization for Efficient Reasoning LLM InferenceYesheng Liang, Haisheng Chen, Song Han, Zhijian LiuICLR 2026 · 被引用 19 次
