Cold Analysis of Rao-Blackwellized Straight-Through Gumbel-Softmax Gradient Estimator
Alexander Shekhovtsov
Abstract
Many problems in machine learning require an estimate of the gradient of an expectation in discrete random variables with respect to the sampling distribution. This work is motivated by the development of the Gumbel-Softmax family of estimators, which use a temperature-controlled relaxation of discrete variables. The state-of-the art in this family, the Gumbel-Rao estimator uses an extra internal sampling to reduce the variance, which may be costly. We analyze this estimator and show that it possesses a zero temperature limit with a surprisingly simple closed form. The limit estimator, called ZGR, has favorable bias and variance properties, it is easy to implement and computationally inexpensive. It decomposes as the average of the straight through (ST) estimator and DARN estimator -two basic but not very well performing on their own estimators. We demonstrate that the simple ST-ZGR family of estimators practically dominates in the biasvariance tradeoffs the whole GR family while also outperforming SOTA unbiased estimators.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6a248915-906c-4928-b1a3-3444a8a293abBuilds on8
- Estimating Gradients for Discrete Random Variables by Sampling without ReplacementWouter Kool, Herke van Hoof, Max WellingICLR 2020 · 59 citations
- Rao-Blackwellizing the Straight-Through Gumbel-Softmax Gradient EstimatorMax B. Paulus, Chris J. Maddison, Andreas KrauseICLR 2021 · 48 citations
- Discrete Representations Strengthen Vision Transformer RobustnessChengzhi Mao, Lu Jiang, Mostafa Dehghani, Carl Vondrick et al.ICLR 2022 · 47 citations
- DisARM: An Antithetic Gradient Estimator for Binary Latent VariablesZhe Dong, Andriy Mnih, George TuckerNeurIPS 2020 · 43 citations
- High-Capacity Expert Binary NetworksAdrian Bulat, Brais Martínez, Georgios TzimiropoulosICLR 2021 · 29 citations
Related papers
- Training Discrete Deep Generative Models via Gapped Straight-Through EstimatorTing-Han Fan, Ta-Chung Chi, Alexander I. Rudnicky, Peter J. RamadgeICML 2022 · 9 citations
- SIMPLE: A Gradient Estimator for k-Subset SamplingKareem Ahmed, Zhe Zeng, Mathias Niepert, Guy Van den BroeckICLR 2023 · 2 citations
- Gradient Estimation with Stochastic Softmax TricksMax B. Paulus, Dami Choi, Daniel Tarlow, Andreas Krause et al.NeurIPS 2020 · 104 citations
- Efficient Marginalization of Discrete and Structured Latent Variables via SparsityGonçalo M. Correia, Vlad Niculae, Wilker Aziz, André F. T. MartinsNeurIPS 2020 · 25 citations
- Leveraging Recursive Gumbel-Max Trick for Approximate Inference in Combinatorial SpacesKirill Struminsky, Artyom Gadetsky, Denis Rakitin, Danil Karpushkin et al.NeurIPS 2021 · 11 citations
