Learning under Quantization for High-Dimensional Linear Regression
Dechen Zhang, Junwei Su, Difan Zou
摘要
The use of low-bit quantization has emerged as an indispensable technique for enabling the efficient training of large-scale models. Despite its widespread empirical success, a rigorous theoretical understanding of its impact on learning performance remains notably absent, even in the simplest linear regression setting. We present the first systematic theoretical study of this fundamental question, analyzing finite-step stochastic gradient descent (SGD) for high-dimensional linear regression under a comprehensive range of quantization targets: data, label, parameter, activation, and gradient. Our novel analytical framework establishes precise algorithm-dependent and data-dependent excess risk bounds that characterize how different quantization affects learning: parameter, activation, and gradient quantization amplify noise during training; data quantization distorts the data spectrum and introduces additional approximation error. Crucially, we distinguish the effects of two quantization schemes: we prove that for additive quantization (with constant quantization steps), the noise amplification benefits from a suppression effect scaled by the batch size, while multiplicative quantization (with input-dependent quantization steps) largely preserves the spectral structure, thereby reducing the spectral distortion. Furthermore, under common polynomial-decay data spectra, we quantitatively compare the risks of multiplicative and additive quantization, drawing a parallel to the comparison between FP and integer quantization methods. Our theory provides a powerful lens to characterize how quantization shapes the learning dynamics of optimization algorithms, paving the way to further explore learning theory under practical hardware constraints.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper13
- Adaptive Gradient Quantization for Data-Parallel SGDFartash Faghri, Iman Tabrizian, Ilia Markov, Dan Alistarh 等NeurIPS 2020 · 被引用 108 次
- A Dynamical Model of Neural Scaling LawsBlake Bordelon, Alexander B. Atanasov, Cengiz PehlevanICML 2024 · 被引用 84 次
- Asynchronous Decentralized SGD with Quantized and Local UpdatesGiorgi Nadiradze, Amirmojtaba Sabour, Peter Davies, Shigang Li 等NeurIPS 2021 · 被引用 61 次
- Last iterate convergence of SGD for Least-Squares in the Interpolation regimeAditya Vardhan Varre, Loucas Pillaud-Vivien, Nicolas FlammarionNeurIPS 2021 · 被引用 52 次
- Tight Nonparametric Convergence Rates for Stochastic Gradient Descent under the Noiseless Linear ModelRaphaël Berthier, Francis R. Bach, Pierre GaillardNeurIPS 2020 · 被引用 49 次
相关 Paper
- Last Iterate Risk Bounds of SGD with Decaying Stepsize for Overparameterized Linear RegressionJingfeng Wu, Difan Zou, Vladimir Braverman, Quanquan Gu 等ICML 2022 · 被引用 38 次
- Risk Bounds of Multi-Pass SGD for Least Squares in the Interpolation RegimeDifan Zou, Jingfeng Wu, Vladimir Braverman, Quanquan Gu 等NeurIPS 2022 · 被引用 9 次
- On the Double Descent of Random Features Models Trained with SGDFanghui Liu, Johan A. K. Suykens, Volkan CevherNeurIPS 2022 · 被引用 11 次
- Risk Bounds of Accelerated SGD for Overparameterized Linear RegressionXuheng Li, Yihe Deng, Jingfeng Wu, Dongruo Zhou 等ICLR 2024 · 被引用 7 次
- Low-Precision Stochastic Gradient Langevin DynamicsRuqi Zhang, Andrew Gordon Wilson, Christopher De SaICML 2022 · 被引用 18 次
