"Give Me BF16 or Give Me Death"? Accuracy-Performance Trade-Offs in LLM Quantization
Eldar Kurtic, Alexandre Noll Marques, Shubhra Pandit, Mark Kurtz, Dan Alistarh
摘要
Quantization is a powerful tool for accelerating large language model (LLM) inference, but the accuracy-performance trade-offs across different formats remain unclear. In this paper, we conduct the most comprehensive empirical study to date, evaluating FP8, INT8, and INT4 quantization across academic benchmarks and real-world tasks on the entire Llama-3.1 model family. Through over 500,000 evaluations, our investigation yields several key findings: (1) FP8 (W8A8-FP) is effectively lossless across all model scales, (2) well-tuned INT8 (W8A8-INT) achieves surprisingly low (1-3%) accuracy degradation, and (3) INT4 weightonly (W4A16-INT) is more competitive than expected, rivaling 8-bit quantization. Further, we investigate the optimal quantization format for different deployments by analyzing inference performance through the popular vLLM framework. Our analysis provides clear deployment recommendations: W4A16 is the most cost-efficient for synchronous setups, while W8A8 dominates in asynchronous continuous batching. For mixed workloads, the optimal choice depends on the specific use case. Our findings offer practical, data-driven guidelines for deploying quantized LLMs at scale-ensuring the best balance between speed, efficiency, and accuracy.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- Bridging the Gap Between Promise and Performance for Microscaling FP4 QuantizationVage Egiazarian, Roberto L. Castro, Denis Kuznedelev, Andrei Panferov 等ICLR 2026 · 被引用 38 次
- The Geometry of LLM Quantization: GPTQ as Babai's Nearest Plane AlgorithmJiale Chen, Yalda Shabanzadeh, Elvir Crnčević, Torsten Hoefler 等ICLR 2026 · 被引用 27 次
- WUSH: Near-Optimal Adaptive Transforms for LLM QuantizationJiale Chen, Vage Egiazarian, Roberto Castro, Torsten Hoefler 等ICML 2026 · 被引用 7 次
- When LLMs get significantly worse: A statistical approach to detect model degradationsJonas M. Kübler, Kailash Budhathoki, Matthäus Kleindessner, Xiong Zhou 等ICLR 2026 · 被引用 6 次
- The Power of Anomaly Detection in Predictive Maintenance: [Experiments & Analysis]Anastasios Papadopoulos, Apostolos Giannoulidis, Anastasios Gounaris, John PaparrizosSIGMOD 2026 · 被引用 4 次
它引用的顶会 Paper19
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- TruthfulQA: Measuring How Models Mimic Human FalsehoodsStephanie Lin, Jacob Hilton, Owain EvansACL 2022 · 被引用 3,228 次
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 被引用 3,037 次
- MiniLM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre-Trained TransformersWenhui Wang, Furu Wei, Li Dong, Hangbo Bao 等NeurIPS 2020 · 被引用 2,727 次
相关 Paper
- Does quantization affect models' performance on long-context tasks?Anmol Mekala, Anirudh Atmakuru, Yixiao Song, Marzena Karpinska 等EMNLP 2025
- COMET: Towards Practical W4A4KV4 LLMs ServingLian Liu, Long Cheng, Haimeng Ren, Zhaohui Xu 等ASPLOS 2025 · 被引用 5 次
- Benchmarking Post-Training Quantization of Large Language Models under Microscaling Floating Point FormatsManyi Zhang, Ji-Fu Li, Zhongao Sun, Haoli Bai 等ACL 2026 · 被引用 2 次
- Understanding Int4 Quantization for Language Models: Latency Speedup, Composability, and Failure CasesXiaoxia Wu, Cheng Li, Reza Yazdani Aminabadi, Zhewei Yao 等ICML 2023 · 被引用 41 次
- Agile-Quant: Activation-Guided Quantization for Faster Inference of LLMs on the EdgeXuan Shen, Peiyan Dong, Lei Lu, Zhenglun Kong 等AAAI 2024 · 被引用 59 次
