High-Dimensional Learning Dynamics of Quantized Models with Straight-Through Estimator
Yuma Ichikawa, Shuhei Kashiwamura, Ayaka Sakata
摘要
Quantized neural network training optimizes a discrete, non-differentiable objective. The straight-through estimator (STE) enables backpropagation through surrogate gradients and is widely used. While previous studies have primarily focused on the properties of surrogate gradients and their convergence, the influence of quantization hyperparameters, such as bit width and quantization range, on learning dynamics remains largely unexplored. We theoretically show that in the high-dimensional limit, STE dynamics converge to a deterministic ordinary differential equation. This reveals that STE training exhibits a plateau followed by a sharp drop in generalization error, with plateau length depending on the quantization range. A fixed-point analysis quantifies the asymptotic deviation from the unquantized linear model. We also extend analytical techniques for stochastic gradient descent to nonlinear transformations of weights and inputs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper14
- Learning curves of generic features maps for realistic datasets with a teacher-student modelBruno Loureiro, Cédric Gerbelot, Hugo Cui, Sebastian Goldt 等NeurIPS 2021 · 被引用 170 次
- PAC-Bayes Compression Bounds So Tight That They Can Explain GeneralizationSanae Lotfi, Marc Finzi, Sanyam Kapoor, Andres Potapczynski 等NeurIPS 2022 · 被引用 98 次
- Phase diagram of Stochastic Gradient Descent in high-dimensional two-layer neural networksRodrigo Veiga, Ludovic Stephan, Bruno Loureiro, Florent Krzakala 等NeurIPS 2022 · 被引用 59 次
- Learning Strides in Convolutional Neural NetworksRachid Riad, Olivier Teboul, David Grangier, Neil ZeghidourICLR 2022 · 被引用 54 次
- Bridging Discrete and Backpropagation: Straight-Through and BeyondLiyuan Liu, Chengyu Dong, Xiaodong Liu, Bin Yu 等NeurIPS 2023 · 被引用 52 次
相关 Paper
- Robust Training of Neural Networks at Arbitrary Precision and SparsityChengxi Ye, Grace Chu, Yanfeng Liu, Yichi Zhang 等ICLR 2026 · 被引用 2 次
- Training Quantised Neural Networks with STE Variants: the Additive Noise Annealing AlgorithmMatteo Spallanzani, Gian Paolo Leonardi, Luca BeniniCVPR 2022 · 被引用 2 次
- Network Quantization With Element-Wise Gradient ScalingJunghyup Lee, Dohyung Kim, Bumsub HamCVPR 2021
- Mixed Precision DNNs: All you need is a good parametrizationStefan Uhlich, Lukas Mauch, Fabien Cardinaux, Kazuki Yoshiyama 等ICLR 2020 · 被引用 159 次
- Cluster-Promoting Quantization with Bit-Drop for Minimizing Network Quantization LossJung Hyun Lee, Jihun Yun, Sung Ju Hwang, Eunho YangICCV 2021 · 被引用 17 次
