Gradient Estimation for Binary Latent Variables via Gradient Variance Clipping
Russell Z. Kunes, Mingzhang Yin, Max Land, Doron Haviv, Dana Pe'er, Simon Tavaré
Abstract
Gradient estimation is often necessary for fitting generative models with discrete latent variables, in contexts such as reinforcement learning and variational autoencoder (VAE) training. The DisARM estimator (Yin et al. 2020; Dong, Mnih, and Tucker 2020) achieves state of the art gradient variance for Bernoulli latent variable models in many contexts. However, DisARM and other estimators have potentially exploding variance near the boundary of the parameter space, where solutions tend to lie. To ameliorate this issue, we propose a new gradient estimator bitflip-1 that has lower variance at the boundaries of the parameter space. As bitflip-1 has complementary properties to existing estimators, we introduce an aggregated estimator, unbiased gradient variance clipping (UGC) that uses either a bitflip-1 or a DisARM gradient update for each coordinate. We theoretically prove that UGC has uniformly lower variance than DisARM. Empirically, we observe that UGC achieves the optimal value of the optimization objectives in toy experiments, discrete VAE training, and in a best subset selection problem.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f26a5e0d-9e36-4f7f-81f5-b9d0a04454e1Cited by top-tier papers3
- High-Dimensional Learning Dynamics of Quantized Models with Straight-Through EstimatorYuma Ichikawa, Shuhei Kashiwamura, Ayaka SakataICML 2026 · 5 citations
- BiPFT: Binary Pre-trained Foundation Transformer with Low-Rank Estimation of Binarization Residual PolynomialsXingrun Xing, Li Du, Xinyuan Wang, Xianlin Zeng et al.AAAI 2024 · 5 citations
- GradInf: Gradient Estimation as Probabilistic InferenceGaurav Arya, Mathieu Huot, Moritz Schauer, Alexander K. Lew et al.PLDI 2026
Builds on4
- Gradient Estimation with Stochastic Softmax TricksMax B. Paulus, Dami Choi, Daniel Tarlow, Andreas Krause et al.NeurIPS 2020 · 104 citations
- DisARM: An Antithetic Gradient Estimator for Binary Latent VariablesZhe Dong, Andriy Mnih, George TuckerNeurIPS 2020 · 43 citations
- Coupled Gradient Estimators for Discrete Latent VariablesZhe Dong, Andriy Mnih, George TuckerNeurIPS 2021 · 14 citations
- ARMS: Antithetic-REINFORCE-Multi-Sample Gradient for Binary VariablesAleksandar Dimitriev, Mingyuan ZhouICML 2021 · 12 citations
Related papers
- Gradient Estimation with Discrete Stein OperatorsJiaxin Shi, Yuhao Zhou, Jessica Hwang, Michalis K. Titsias et al.NeurIPS 2022 · 27 citations
- Estimating Gradients for Discrete Random Variables by Sampling without ReplacementWouter Kool, Herke van Hoof, Max WellingICLR 2020 · 59 citations
- Training Discrete Deep Generative Models via Gapped Straight-Through EstimatorTing-Han Fan, Ta-Chung Chi, Alexander I. Rudnicky, Peter J. RamadgeICML 2022 · 9 citations
- Hindsight Network Credit Assignment: Efficient Credit Assignment in Networks of Discrete Stochastic UnitsKenny YoungAAAI 2022
- CARMS: Categorical-Antithetic-REINFORCE Multi-Sample Gradient EstimatorAlek Dimitriev, Mingyuan ZhouNeurIPS 2021 · 10 citations
