High-Dimensional Learning Dynamics of Quantized Models with Straight-Through Estimator
Yuma Ichikawa, Shuhei Kashiwamura, Ayaka Sakata
Abstract
Quantized neural network training optimizes a discrete, non-differentiable objective. The straight-through estimator (STE) enables backpropagation through surrogate gradients and is widely used. While previous studies have primarily focused on the properties of surrogate gradients and their convergence, the influence of quantization hyperparameters, such as bit width and quantization range, on learning dynamics remains largely unexplored. We theoretically show that in the high-dimensional limit, STE dynamics converge to a deterministic ordinary differential equation. This reveals that STE training exhibits a plateau followed by a sharp drop in generalization error, with plateau length depending on the quantization range. A fixed-point analysis quantifies the asymptotic deviation from the unquantized linear model. We also extend analytical techniques for stochastic gradient descent to nonlinear transformations of weights and inputs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on14
- Learning curves of generic features maps for realistic datasets with a teacher-student modelBruno Loureiro, Cédric Gerbelot, Hugo Cui, Sebastian Goldt et al.NeurIPS 2021 · 170 citations
- PAC-Bayes Compression Bounds So Tight That They Can Explain GeneralizationSanae Lotfi, Marc Finzi, Sanyam Kapoor, Andres Potapczynski et al.NeurIPS 2022 · 98 citations
- Phase diagram of Stochastic Gradient Descent in high-dimensional two-layer neural networksRodrigo Veiga, Ludovic Stephan, Bruno Loureiro, Florent Krzakala et al.NeurIPS 2022 · 59 citations
- Learning Strides in Convolutional Neural NetworksRachid Riad, Olivier Teboul, David Grangier, Neil ZeghidourICLR 2022 · 54 citations
- Bridging Discrete and Backpropagation: Straight-Through and BeyondLiyuan Liu, Chengyu Dong, Xiaodong Liu, Bin Yu et al.NeurIPS 2023 · 52 citations
Related papers
- Robust Training of Neural Networks at Arbitrary Precision and SparsityChengxi Ye, Grace Chu, Yanfeng Liu, Yichi Zhang et al.ICLR 2026 · 2 citations
- Training Quantised Neural Networks with STE Variants: the Additive Noise Annealing AlgorithmMatteo Spallanzani, Gian Paolo Leonardi, Luca BeniniCVPR 2022 · 2 citations
- Network Quantization With Element-Wise Gradient ScalingJunghyup Lee, Dohyung Kim, Bumsub HamCVPR 2021
- Mixed Precision DNNs: All you need is a good parametrizationStefan Uhlich, Lukas Mauch, Fabien Cardinaux, Kazuki Yoshiyama et al.ICLR 2020 · 159 citations
- Cluster-Promoting Quantization with Bit-Drop for Minimizing Network Quantization LossJung Hyun Lee, Jihun Yun, Sung Ju Hwang, Eunho YangICCV 2021 · 17 citations
