Scaling Laws for Floating-Point Quantization Training
Xingwu Sun, Shuaipeng Li, Ruobing Xie, Weidong Han, Kan Wu, Zhen Yang, Yixing Li, An Wang, Shuai Li, Jinbao Xue, Yu Cheng, Yangyu Tao
Abstract
Low-precision training is considered an effective strategy for reducing both training and downstream inference costs. Previous scaling laws for precision mainly focus on integer quantization, which pay less attention to the constituents in floating-point (FP) quantization, and thus cannot well fit the LLM losses in this scenario. In contrast, while FP quantization training is more commonly implemented in production, it's research has been relatively superficial. In this paper, we thoroughly explore the effects of FP quantization targets, exponent bits, mantissa bits, and the calculation granularity of the scaling factor in FP quantization training performance of LLM models. In addition to an accurate FP quantization unified scaling law, we also provide valuable suggestions for the community: (1) Exponent bits contribute slightly more to the model performance than mantissa bits. We provide the optimal exponent-mantissa bit ratio for different bit numbers, which is available for future reference by hardware manufacturers; (2) We discover the formation of the critical data size in low-precision LLM training. Too much training data exceeding the critical data size will inversely bring in degradation of LLM performance; (3) The optimal FP quantization precision is directly proportional to the computational power, but within a wide computational power range. We estimate that the best cost-performance precision should lie between 4-8 bits.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1ec3f904-bb3f-41f5-bad7-ff313416582fCited by top-tier papers5
- Scaling Law for Quantization-Aware TrainingMengzhao Chen, Chaoyi Zhang, Jing Liu, Zeng et al.ICML 2026 · 16 citations
- Unified Scaling Laws for Compressed RepresentationsAndrei Panferov, Alexandra Volkova, Ionut-Vlad Modoranu, Vage Egiazarian et al.NeurIPS 2025 · 5 citations
- Scaling Laws for Precision in High-Dimensional Linear RegressionDechen Zhang, Xuan Tang, Yingyu Liang, Difan ZouICML 2026 · 2 citations
- Learning under Quantization for High-Dimensional Linear RegressionDechen Zhang, Junwei Su, Difan ZouICLR 2026 · 1 citation
- How Do Large Language Monkeys Get Their Power (Laws)?Rylan Schaeffer, Joshua Kazdan, John Hughes, Jordan Juravsky et al.ICML 2025
Builds on8
- Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space DualityTri Dao, Albert GuICML 2024 · 1,407 citations
- The case for 4-bit precision: k-bit Inference Scaling LawsTim Dettmers, Luke ZettlemoyerICML 2023 · 315 citations
- Extreme Compression of Large Language Models via Additive QuantizationVage Egiazarian, Andrei Panferov, Denis Kuznedelev, Elias Frantar et al.ICML 2024 · 187 citations
- FP8 Quantization: The Power of the ExponentAndrey Kuzmin, Mart van Baalen, Yuwei Ren, Markus Nagel et al.NeurIPS 2022 · 154 citations
- Scaling Laws in Linear Regression: Compute, Parameters, and DataLicong Lin, Jingfeng Wu, Sham M. Kakade, Peter L. Bartlett et al.NeurIPS 2024 · 57 citations
Related papers
- Scaling Laws for PrecisionTanishq Kumar, Zachary Ankner, Benjamin Frederick Spector, Blake Bordelon et al.ICLR 2025
- Low-Bit Quantization Favors Undertrained LLMsXu Ouyang, Tao Ge, Thomas Hartvigsen, Zhisong Zhang et al.ACL 2025 · 3 citations
- LLM-FP4: 4-Bit Floating-Point Quantized TransformersShih-Yang Liu, Zechun Liu, Xijie Huang, Pingcheng Dong et al.EMNLP 2023 · 34 citations
- Benchmarking Post-Training Quantization of Large Language Models under Microscaling Floating Point FormatsManyi Zhang, Ji-Fu Li, Zhongao Sun, Haoli Bai et al.ACL 2026 · 2 citations
- Compute-Optimal Quantization-Aware TrainingAleksandr Dremov, David Grangier, Angelos Katharopoulos, Awni HannunICLR 2026 · 4 citations
