Gradient Estimation for Binary Latent Variables via Gradient Variance Clipping
Russell Z. Kunes, Mingzhang Yin, Max Land, Doron Haviv, Dana Pe'er, Simon Tavaré
摘要
Gradient estimation is often necessary for fitting generative models with discrete latent variables, in contexts such as reinforcement learning and variational autoencoder (VAE) training. The DisARM estimator (Yin et al. 2020; Dong, Mnih, and Tucker 2020) achieves state of the art gradient variance for Bernoulli latent variable models in many contexts. However, DisARM and other estimators have potentially exploding variance near the boundary of the parameter space, where solutions tend to lie. To ameliorate this issue, we propose a new gradient estimator bitflip-1 that has lower variance at the boundaries of the parameter space. As bitflip-1 has complementary properties to existing estimators, we introduce an aggregated estimator, unbiased gradient variance clipping (UGC) that uses either a bitflip-1 or a DisARM gradient update for each coordinate. We theoretically prove that UGC has uniformly lower variance than DisARM. Empirically, we observe that UGC achieves the optimal value of the optimization objectives in toy experiments, discrete VAE training, and in a best subset selection problem.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- High-Dimensional Learning Dynamics of Quantized Models with Straight-Through EstimatorYuma Ichikawa, Shuhei Kashiwamura, Ayaka SakataICML 2026 · 被引用 5 次
- BiPFT: Binary Pre-trained Foundation Transformer with Low-Rank Estimation of Binarization Residual PolynomialsXingrun Xing, Li Du, Xinyuan Wang, Xianlin Zeng 等AAAI 2024 · 被引用 5 次
- GradInf: Gradient Estimation as Probabilistic InferenceGaurav Arya, Mathieu Huot, Moritz Schauer, Alexander K. Lew 等PLDI 2026
它引用的顶会 Paper4
- Gradient Estimation with Stochastic Softmax TricksMax B. Paulus, Dami Choi, Daniel Tarlow, Andreas Krause 等NeurIPS 2020 · 被引用 104 次
- DisARM: An Antithetic Gradient Estimator for Binary Latent VariablesZhe Dong, Andriy Mnih, George TuckerNeurIPS 2020 · 被引用 43 次
- Coupled Gradient Estimators for Discrete Latent VariablesZhe Dong, Andriy Mnih, George TuckerNeurIPS 2021 · 被引用 14 次
- ARMS: Antithetic-REINFORCE-Multi-Sample Gradient for Binary VariablesAleksandar Dimitriev, Mingyuan ZhouICML 2021 · 被引用 12 次
相关 Paper
- Gradient Estimation with Discrete Stein OperatorsJiaxin Shi, Yuhao Zhou, Jessica Hwang, Michalis K. Titsias 等NeurIPS 2022 · 被引用 27 次
- Estimating Gradients for Discrete Random Variables by Sampling without ReplacementWouter Kool, Herke van Hoof, Max WellingICLR 2020 · 被引用 59 次
- Training Discrete Deep Generative Models via Gapped Straight-Through EstimatorTing-Han Fan, Ta-Chung Chi, Alexander I. Rudnicky, Peter J. RamadgeICML 2022 · 被引用 9 次
- Hindsight Network Credit Assignment: Efficient Credit Assignment in Networks of Discrete Stochastic UnitsKenny YoungAAAI 2022
- CARMS: Categorical-Antithetic-REINFORCE Multi-Sample Gradient EstimatorAlek Dimitriev, Mingyuan ZhouNeurIPS 2021 · 被引用 10 次
