Cold Analysis of Rao-Blackwellized Straight-Through Gumbel-Softmax Gradient Estimator
Alexander Shekhovtsov
摘要
Many problems in machine learning require an estimate of the gradient of an expectation in discrete random variables with respect to the sampling distribution. This work is motivated by the development of the Gumbel-Softmax family of estimators, which use a temperature-controlled relaxation of discrete variables. The state-of-the art in this family, the Gumbel-Rao estimator uses an extra internal sampling to reduce the variance, which may be costly. We analyze this estimator and show that it possesses a zero temperature limit with a surprisingly simple closed form. The limit estimator, called ZGR, has favorable bias and variance properties, it is easy to implement and computationally inexpensive. It decomposes as the average of the straight through (ST) estimator and DARN estimator -two basic but not very well performing on their own estimators. We demonstrate that the simple ST-ZGR family of estimators practically dominates in the biasvariance tradeoffs the whole GR family while also outperforming SOTA unbiased estimators.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper8
- Estimating Gradients for Discrete Random Variables by Sampling without ReplacementWouter Kool, Herke van Hoof, Max WellingICLR 2020 · 被引用 59 次
- Rao-Blackwellizing the Straight-Through Gumbel-Softmax Gradient EstimatorMax B. Paulus, Chris J. Maddison, Andreas KrauseICLR 2021 · 被引用 48 次
- Discrete Representations Strengthen Vision Transformer RobustnessChengzhi Mao, Lu Jiang, Mostafa Dehghani, Carl Vondrick 等ICLR 2022 · 被引用 47 次
- DisARM: An Antithetic Gradient Estimator for Binary Latent VariablesZhe Dong, Andriy Mnih, George TuckerNeurIPS 2020 · 被引用 43 次
- High-Capacity Expert Binary NetworksAdrian Bulat, Brais Martínez, Georgios TzimiropoulosICLR 2021 · 被引用 29 次
相关 Paper
- Training Discrete Deep Generative Models via Gapped Straight-Through EstimatorTing-Han Fan, Ta-Chung Chi, Alexander I. Rudnicky, Peter J. RamadgeICML 2022 · 被引用 9 次
- SIMPLE: A Gradient Estimator for k-Subset SamplingKareem Ahmed, Zhe Zeng, Mathias Niepert, Guy Van den BroeckICLR 2023 · 被引用 2 次
- Gradient Estimation with Stochastic Softmax TricksMax B. Paulus, Dami Choi, Daniel Tarlow, Andreas Krause 等NeurIPS 2020 · 被引用 104 次
- Efficient Marginalization of Discrete and Structured Latent Variables via SparsityGonçalo M. Correia, Vlad Niculae, Wilker Aziz, André F. T. MartinsNeurIPS 2020 · 被引用 25 次
- Leveraging Recursive Gumbel-Max Trick for Approximate Inference in Combinatorial SpacesKirill Struminsky, Artyom Gadetsky, Denis Rakitin, Danil Karpushkin 等NeurIPS 2021 · 被引用 11 次
