Scaling Laws for Precision in High-Dimensional Linear Regression
Dechen Zhang, Xuan Tang, Yingyu Liang, Difan Zou
Abstract
Low-precision training is critical for optimizing the trade-off between model quality and training costs, necessitating the joint allocation of model size, dataset size, and numerical precision. While empirical scaling laws suggest that quantization impacts effective model and data capacities or acts as an additive error, the theoretical mechanisms governing these effects remain largely unexplored. In this work, we initiate a theoretical study of scaling laws for low-precision training within a high-dimensional sketched linear regression framework. By analyzing multiplicative (signal-dependent) and additive (signal-independent) quantization, we identify a critical dichotomy in their scaling behaviors. Our analysis reveals that in the worst case, while both schemes introduce an additive error and degrade the effective data size, they exhibit distinct effects on effective model size: multiplicative quantization maintains the full-precision model size, whereas additive quantization reduces the effective model size. Numerical experiments validate our theoretical findings. By rigorously characterizing the complex interplay among model scale, dataset size, and quantization error, our work provides a principled theoretical basis for optimizing training protocols under practical hardware constraints.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5f8dc3a9-be14-44fb-b2ee-5998df2e102dBuilds on17
- The case for 4-bit precision: k-bit Inference Scaling LawsTim Dettmers, Luke ZettlemoyerICML 2023 · 315 citations
- Ultra-Low Precision 4-bit Training of Deep Neural NetworksXiao Sun, Naigang Wang, Chia-Yu Chen, Jiamin Ni et al.NeurIPS 2020 · 227 citations
- Scaling Laws with Vocabulary: Larger Models Deserve Larger VocabulariesChaofan Tao, Qian Liu, Longxu Dou, Niklas Muennighoff et al.NeurIPS 2024 · 135 citations
- Stable and low-precision training for large-scale vision-language modelsMitchell Wortsman, Tim Dettmers, Luke Zettlemoyer, Ari Morcos et al.NeurIPS 2023 · 101 citations
- A Dynamical Model of Neural Scaling LawsBlake Bordelon, Alexander B. Atanasov, Cengiz PehlevanICML 2024 · 84 citations
Related papers
- Learning under Quantization for High-Dimensional Linear RegressionDechen Zhang, Junwei Su, Difan ZouICLR 2026 · 1 citation
- Scaling Laws for PrecisionTanishq Kumar, Zachary Ankner, Benjamin Frederick Spector, Blake Bordelon et al.ICLR 2025
- Scaling Laws for Floating-Point Quantization TrainingXingwu Sun, Shuaipeng Li, Ruobing Xie, Weidong Han et al.ICML 2025
- M+Adam: Low-Precision Training via Additive–Multiplicative OptimizationXiaoyuan Liang, Sebastian Loeschcke, Mads Toftrup, Anima AnandkumarICML 2026 · 1 citation
- Low-Bit Quantization Favors Undertrained LLMsXu Ouyang, Tao Ge, Thomas Hartvigsen, Zhisong Zhang et al.ACL 2025 · 3 citations
