DisARM: An Antithetic Gradient Estimator for Binary Latent Variables
Zhe Dong, Andriy Mnih, George Tucker
Abstract
Training models with discrete latent variables is challenging due to the difficulty of estimating the gradients accurately. Much of the recent progress has been achieved by taking advantage of continuous relaxations of the system, which are not always available or even possible. The Augment-REINFORCE-Merge (ARM) estimator provides an alternative that, instead of relaxation, uses continuous augmentation. Applying antithetic sampling over the augmenting variables yields a relatively low-variance and unbiased estimator applicable to any model with binary latent variables. However, while antithetic sampling reduces variance, the augmentation process increases variance. We show that ARM can be improved by analytically integrating out the randomness introduced by the augmentation process, guaranteeing substantial variance reduction. Our estimator, DisARM, is simple to implement and has the same computational cost as ARM. We evaluate DisARM on several generative modeling benchmarks and show that it consistently outperforms ARM and a strong independent sample baseline in terms of both variance and log-likelihood. Furthermore, we propose a local version of DisARM designed for optimizing the multi-sample variational bound, and show that it outperforms VIMCO, the current state-of-the-art method. Code and additional information: https://sites.google.com/view/disarm-estimator .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 52d20e9d-e6a3-4a88-8745-736e8d5d4420Cited by top-tier papers19
- VarGrad: A Low-Variance Gradient Estimator for Variational InferenceLorenz Richter, Ayman Boustati, Nikolas Nüsken, Francisco J. R. Ruiz et al.NeurIPS 2020 · 90 citations
- Structured Graph Convolutional Networks with Stochastic Masks for Recommender SystemsHuiyuan Chen, Lan Wang, Yusan Lin, Chin-Chia Michael Yeh et al.SIGIR 2021 · 60 citations
- Bridging Discrete and Backpropagation: Straight-Through and BeyondLiyuan Liu, Chengyu Dong, Xiaodong Liu, Bin Yu et al.NeurIPS 2023 · 52 citations
- Knowledge-refined Denoising Network for Robust RecommendationXinjun Zhu, Yuntao Du, Yuren Mao, Lu Chen et al.SIGIR 2023 · 36 citations
- Contextual Dropout: An Efficient Sample-Dependent Dropout ModuleXinjie Fan, Shujian Zhang, Korawat Tanwisuth, Xiaoning Qian et al.ICLR 2021 · 34 citations
Related papers
- ARMS: Antithetic-REINFORCE-Multi-Sample Gradient for Binary VariablesAleksandar Dimitriev, Mingyuan ZhouICML 2021 · 12 citations
- Coupled Gradient Estimators for Discrete Latent VariablesZhe Dong, Andriy Mnih, George TuckerNeurIPS 2021 · 14 citations
- CARMS: Categorical-Antithetic-REINFORCE Multi-Sample Gradient EstimatorAlek Dimitriev, Mingyuan ZhouNeurIPS 2021 · 10 citations
- Gradient Estimation for Binary Latent Variables via Gradient Variance ClippingRussell Z. Kunes, Mingzhang Yin, Max Land, Doron Haviv et al.AAAI 2023 · 5 citations
- Gradient Estimation with Discrete Stein OperatorsJiaxin Shi, Yuhao Zhou, Jessica Hwang, Michalis K. Titsias et al.NeurIPS 2022 · 27 citations
