Training Discrete Deep Generative Models via Gapped Straight-Through Estimator
Ting-Han Fan, Ta-Chung Chi, Alexander I. Rudnicky, Peter J. Ramadge
Abstract
While deep generative models have succeeded in image processing, natural language processing, and reinforcement learning, training that involves discrete random variables remains challenging due to the high variance of its gradient estimation process. Monte Carlo is a common solution used in most variance reduction approaches. However, this involves time-consuming resampling and multiple function evaluations. We propose a Gapped Straight-Through (GST) estimator to reduce the variance without incurring resampling overhead. This estimator is inspired by the essential properties of Straight-Through Gumbel-Softmax. We determine these properties and show via an ablation study that they are essential. Experiments demonstrate that the proposed GST estimator enjoys better performance compared to strong baselines on two discrete deep generative modeling tasks, MNIST-VAE and ListOps.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 21425b2d-d36b-4933-9120-f5cd8f53e774Cited by top-tier papers1
Ask how each one uses itBuilds on6
- Score-Based Generative Modeling through Stochastic Differential EquationsYang Song, Jascha Sohl-Dickstein, Diederik P. Kingma, Abhishek Kumar et al.ICLR 2021 · 1,270 citations
- Discretizing Continuous Action Space for On-Policy OptimizationYunhao Tang, Shipra AgrawalAAAI 2020 · 150 citations
- Gradient Estimation with Stochastic Softmax TricksMax B. Paulus, Dami Choi, Daniel Tarlow, Andreas Krause et al.NeurIPS 2020 · 104 citations
- Rao-Blackwellizing the Straight-Through Gumbel-Softmax Gradient EstimatorMax B. Paulus, Chris J. Maddison, Andreas KrauseICLR 2021 · 48 citations
- DisARM: An Antithetic Gradient Estimator for Binary Latent VariablesZhe Dong, Andriy Mnih, George TuckerNeurIPS 2020 · 43 citations
Related papers
- Cold Analysis of Rao-Blackwellized Straight-Through Gumbel-Softmax Gradient EstimatorAlexander ShekhovtsovICML 2023 · 2 citations
- Low Bias Low Variance Gradient Estimates for Boolean Stochastic NetworksAdeel Pervez, Taco Cohen, Efstratios GavvesICML 2020 · 10 citations
- Discrete Variational Autoencoding via Policy SearchMichael Drolet, Firas Al-Hafez, Aditya Bhatt, Jan Peters et al.ICLR 2026
- Estimating Gradients for Discrete Random Variables by Sampling without ReplacementWouter Kool, Herke van Hoof, Max WellingICLR 2020 · 59 citations
- Gradient Estimation with Discrete Stein OperatorsJiaxin Shi, Yuhao Zhou, Jessica Hwang, Michalis K. Titsias et al.NeurIPS 2022 · 27 citations
