Low Bias Low Variance Gradient Estimates for Boolean Stochastic Networks
Adeel Pervez, Taco Cohen, Efstratios Gavves
Abstract
Stochastic neural networks with discrete random variables are an important class of models for their expressiveness and interpretability. Since direct differentiation and backpropagation is not possible, Monte Carlo gradient estimation techniques are a popular alternative. Efficient stochastic gradient estimators, such Straight-Through and Gumbel-Softmax, work well for shallow stochastic models. Their performance, however, suffers with hierarchical, more complex models. We focus on stochastic networks with Boolean latent variables. To analyze such networks, we introduce the framework of harmonic analysis for Boolean functions to derive an analytic formulation for the bias and variance in the Straight-Through estimator. Exploiting these formulations, we propose FouST, a low-bias and low-variance gradient estimation algorithm that is just as efficient. Extensive experiments show that FouST performs favorably compared to state-of-the-art biased estimators and is much faster than unbiased ones. 1 Training a nonlinear sigmoid belief network model on GPU with two stochastic layers on MNIST with REBAR took 1.5 days.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 17265f16-fe01-4faf-8cf8-35191a8a6b8eCited by top-tier papers5
- Multi-Facet Clustering Variational AutoencodersFabian Falck, Haoting Zhang, Matthew Willetts, George Nicholson et al.NeurIPS 2021 · 57 citations
- Bridging Discrete and Backpropagation: Straight-Through and BeyondLiyuan Liu, Chengyu Dong, Xiaodong Liu, Bin Yu et al.NeurIPS 2023 · 52 citations
- Spectral Smoothing Unveils Phase Transitions in Hierarchical Variational AutoencodersAdeel Pervez, Efstratios GavvesICML 2021 · 4 citations
- Efficient Learning of Discrete-Continuous Computation GraphsDavid Friede, Mathias NiepertNeurIPS 2021 · 3 citations
- Cold Analysis of Rao-Blackwellized Straight-Through Gumbel-Softmax Gradient EstimatorAlexander ShekhovtsovICML 2023 · 2 citations
Related papers
- Training Discrete Deep Generative Models via Gapped Straight-Through EstimatorTing-Han Fan, Ta-Chung Chi, Alexander I. Rudnicky, Peter J. RamadgeICML 2022 · 9 citations
- Rao-Blackwellizing the Straight-Through Gumbel-Softmax Gradient EstimatorMax B. Paulus, Chris J. Maddison, Andreas KrauseICLR 2021 · 48 citations
- Path Sample-Analytic Gradient Estimators for Stochastic Binary NetworksAlexander Shekhovtsov, Viktor Yanush, Boris FlachNeurIPS 2020 · 14 citations
- Gradient Estimation with Stochastic Softmax TricksMax B. Paulus, Dami Choi, Daniel Tarlow, Andreas Krause et al.NeurIPS 2020 · 104 citations
- Marginalized Stochastic Natural Gradients for Black-Box Variational InferenceGeng Ji, Debora Sujono, Erik B. SudderthICML 2021 · 9 citations
