Bridging Discrete and Backpropagation: Straight-Through and Beyond
Liyuan Liu, Chengyu Dong, Xiaodong Liu, Bin Yu, Jianfeng Gao
Abstract
Backpropagation, the cornerstone of deep learning, is limited to computing gradients for continuous variables. This limitation poses challenges for problems involving discrete latent variables. To address this issue, we propose a novel approach to approximate the gradient of parameters involved in generating discrete latent variables. First, we examine the widely used Straight-Through (ST) heuristic and demonstrate that it works as a first-order approximation of the gradient. Guided by our findings, we propose ReinMax, which achieves second-order accuracy by integrating Heun's method, a second-order numerical method for solving ODEs. ReinMax does not require Hessian or other second-order derivatives, thus having negligible computation overheads. Extensive experimental results on various tasks demonstrate the superiority of ReinMax over the state of the art. Implementations are released at https://github.com/microsoft/ReinMax.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a6b52dd5-120f-4931-b614-519d6b341fb6Cited by top-tier papers17
- DISP-LLM: Dimension-Independent Structural Pruning for Large Language ModelsShangqian Gao, Chi-Heng Lin, Ting Hua, Zheng Tang et al.NeurIPS 2024 · 42 citations
- MDNS: Masked Diffusion Neural Sampler via Stochastic Optimal ControlYuchen Zhu, Wei Guo, Jaemoo Choi, Guan-Horng Liu et al.NeurIPS 2025 · 24 citations
- FedBAT: Communication-Efficient Federated Learning via Learnable BinarizationShiwei Li, Wenchao Xu, Haozhao Wang, Xing Tang et al.ICML 2024 · 13 citations
- Masked Random Noise for Communication-Efficient Federated LearningShiwei Li, Yingyi Cheng, Haozhao Wang, Xing Tang et al.ACM MM 2024 · 7 citations
- High-Dimensional Learning Dynamics of Quantized Models with Straight-Through EstimatorYuma Ichikawa, Shuhei Kashiwamura, Ayaka SakataICML 2026 · 5 citations
Builds on7
- On the Variance of the Adaptive Learning Rate and BeyondLiyuan Liu, Haoming Jiang, Pengcheng He, Weizhu Chen et al.ICLR 2020 · 2,210 citations
- Rao-Blackwellizing the Straight-Through Gumbel-Softmax Gradient EstimatorMax B. Paulus, Chris J. Maddison, Andreas KrauseICLR 2021 · 48 citations
- DisARM: An Antithetic Gradient Estimator for Binary Latent VariablesZhe Dong, Andriy Mnih, George TuckerNeurIPS 2020 · 43 citations
- Gradient Estimation with Discrete Stein OperatorsJiaxin Shi, Yuhao Zhou, Jessica Hwang, Michalis K. Titsias et al.NeurIPS 2022 · 27 citations
- Coupled Gradient Estimators for Discrete Latent VariablesZhe Dong, Andriy Mnih, George TuckerNeurIPS 2021 · 14 citations
Related papers
- Do Residual Neural Networks discretize Neural Ordinary Differential Equations?Michael E. Sander, Pierre Ablin, Gabriel PeyréNeurIPS 2022 · 42 citations
- Storchastic: A Framework for General Stochastic Automatic DifferentiationEmile van Krieken, Jakub M. Tomczak, Annette ten TeijeNeurIPS 2021 · 19 citations
- Second-Order Neural ODE OptimizerGuan-Horng Liu, Tianrong Chen, Evangelos A. TheodorouNeurIPS 2021 · 20 citations
- Categorical Reparameterization with Denoising Diffusion ModelsSamson Gourevitch, Alain Oliviero Durmus, Eric Moulines, Jimmy Olsson et al.ICML 2026 · 1 citation
- Second-order forward-mode optimization of recurrent neural networks for neuroscienceYoujing Yu, Rui Xia, Qingxi Ma, Máté Lengyel et al.NeurIPS 2024 · 6 citations
