Improving Block-Wise LLM Quantization by 4-bit Block-Wise Optimal Float (BOF4): Analysis and Variations
Patrick Blumenberg, Thomas Graave, Tim Fingscheidt
Abstract
Large language models (LLMs) demand extensive memory capacity during both fine-tuning and inference. To enable memory-efficient fine-tuning, existing methods apply block-wise quantization techniques, such as NF4 and AF4, to the network weights. We show that these quantization techniques incur suboptimal quantization errors. Therefore, as a first novelty, we propose an optimization approach for blockwise quantization. Using this method, we design a family of quantizers named 4-bit block-wise optimal float (BOF4), which consistently reduces the quantization error compared to both baseline methods. We provide both a theoretical and a datadriven solution for the optimization process and prove their practical equivalence. Secondly, we propose a modification to the employed normalization method based on the signed absolute block maximum (BOF4-S), enabling further reduction of the quantization error and empirically achieving less degradation in language modeling performance. Thirdly, we explore additional variations of block-wise quantization methods applied to LLMs through an experimental study on the importance of accurately representing zero and large-amplitude weights on the one hand, and optimization towards various error metrics on the other hand. Lastly, we introduce a mixed-precision quantization strategy dubbed outlier-preserving quantization (OPQ) to address the distributional mismatch induced by outlier weights in block-wise quantization. By storing outlier weights in 16-bit precision (OPQ) while applying BOF4-S, we achieve top performance among 4-bit blockwise quantization techniques w.r.t. perplexity. Preprint. Under review.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 13734fde-9da4-45f8-a980-fe47816c76f3Cited by top-tier papers1
Ask how each one uses itBuilds on12
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 5,863 citations
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 3,037 citations
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao et al.AAAI 2020 · 2,916 citations
Related papers
- OWQ: Outlier-Aware Weight Quantization for Efficient Fine-Tuning and Inference of Large Language ModelsChanghun Lee, Jungyu Jin, Taesu Kim, Hyungjun Kim et al.AAAI 2024 · 134 citations
- RoMeo: Mitigating Dual-dimensional Outliers with Rotated Mixed Precision QuantizationQihao Zhang, Mingliang Tang, Mingshu Zhai, Kinman Lei et al.PPoPP 2026
- BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM InferenceWonsuk Jang, Thierry TambeICML 2025
- BiE: Bi-Exponent Block Floating-Point for Large Language Models QuantizationLancheng Zou, Wenqian Zhao, Shuo Yin, Chen Bai et al.ICML 2024 · 13 citations
- Revisiting Block-based Quantisation: What is Important for Sub-8-bit LLM Inference?Cheng Zhang, Jianyi Cheng, Ilia Shumailov, George A. Constantinides et al.EMNLP 2023 · 5 citations
