Gradient Estimation with Discrete Stein Operators
Jiaxin Shi, Yuhao Zhou, Jessica Hwang, Michalis K. Titsias, Lester Mackey
Abstract
Gradient estimation -- approximating the gradient of an expectation with respect to the parameters of a distribution -- is central to the solution of many machine learning problems. However, when the distribution is discrete, most common gradient estimators suffer from excessive variance. To improve the quality of gradient estimation, we introduce a variance reduction technique based on Stein operators for discrete distributions. We then use this technique to build flexible control variates for the REINFORCE leave-one-out estimator. Our control variates can be adapted online to minimize variance and do not require extra evaluations of the target function. In benchmark generative modeling tasks such as training binary variational autoencoders, our gradient estimator achieves substantially lower variance than state-of-the-art estimators with the same number of function evaluations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f101e026-d53e-49bb-b82d-3bd948824f8eCited by top-tier papers12
- Simplified and Generalized Masked Diffusion for Discrete DataJiaxin Shi, Kehang Han, Zhe Wang, Arnaud Doucet et al.NeurIPS 2024 · 693 citations
- Informed Correctors for Discrete Diffusion ModelsYixiu Zhao, Jiaxin Shi, Feng Chen, Shaul Druckmann et al.NeurIPS 2025 · 61 citations
- Bridging Discrete and Backpropagation: Straight-Through and BeyondLiyuan Liu, Chengyu Dong, Xiaodong Liu, Bin Yu et al.NeurIPS 2023 · 52 citations
- Differentiable Sampling of Categorical Distributions Using the CatLog-Derivative TrickLennert De Smet, Emanuele Sansone, Pedro Zuidberg Dos MartiresNeurIPS 2023 · 17 citations
- Discrete Neural Flow Samplers with Locally Equivariant TransformerZijing Ou, Ruixiang Zhang, Yingzhen LiNeurIPS 2025 · 14 citations
Builds on8
- NVAE: A Deep Hierarchical Variational AutoencoderArash Vahdat, Jan KautzNeurIPS 2020 · 1,141 citations
- Implicit MLE: Backpropagating Through Discrete Exponential Family DistributionsMathias Niepert, Pasquale Minervini, Luca FranceschiNeurIPS 2021 · 121 citations
- Oops I Took A Gradient: Scalable Sampling for Discrete DistributionsWill Grathwohl, Kevin Swersky, Milad Hashemi, David Duvenaud et al.ICML 2021 · 113 citations
- VarGrad: A Low-Variance Gradient Estimator for Variational InferenceLorenz Richter, Ayman Boustati, Nikolas Nüsken, Francisco J. R. Ruiz et al.NeurIPS 2020 · 90 citations
- Estimating Gradients for Discrete Random Variables by Sampling without ReplacementWouter Kool, Herke van Hoof, Max WellingICLR 2020 · 59 citations
Related papers
- Low-Variance Black-Box Gradient Estimates for the Plackett-Luce DistributionArtyom Gadetsky, Kirill Struminsky, Christopher Robinson, Novi Quadrianto et al.AAAI 2020 · 11 citations
- Hindsight Network Credit Assignment: Efficient Credit Assignment in Networks of Discrete Stochastic UnitsKenny YoungAAAI 2022
- Multilevel Control FunctionalKaiyu Li, Yiming Yang, Xiaoyuan Cheng, Yi He et al.ICLR 2026 · 1 citation
- Training Discrete Deep Generative Models via Gapped Straight-Through EstimatorTing-Han Fan, Ta-Chung Chi, Alexander I. Rudnicky, Peter J. RamadgeICML 2022 · 9 citations
- Approximation Based Variance Reduction for Reparameterization GradientsTomas Geffner, Justin DomkeNeurIPS 2020 · 13 citations
