Low Bias Low Variance Gradient Estimates for Boolean Stochastic Networks
Adeel Pervez, Taco Cohen, Efstratios Gavves
摘要
Stochastic neural networks with discrete random variables are an important class of models for their expressiveness and interpretability. Since direct differentiation and backpropagation is not possible, Monte Carlo gradient estimation techniques are a popular alternative. Efficient stochastic gradient estimators, such Straight-Through and Gumbel-Softmax, work well for shallow stochastic models. Their performance, however, suffers with hierarchical, more complex models. We focus on stochastic networks with Boolean latent variables. To analyze such networks, we introduce the framework of harmonic analysis for Boolean functions to derive an analytic formulation for the bias and variance in the Straight-Through estimator. Exploiting these formulations, we propose FouST, a low-bias and low-variance gradient estimation algorithm that is just as efficient. Extensive experiments show that FouST performs favorably compared to state-of-the-art biased estimators and is much faster than unbiased ones. 1 Training a nonlinear sigmoid belief network model on GPU with two stochastic layers on MNIST with REBAR took 1.5 days.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Multi-Facet Clustering Variational AutoencodersFabian Falck, Haoting Zhang, Matthew Willetts, George Nicholson 等NeurIPS 2021 · 被引用 57 次
- Bridging Discrete and Backpropagation: Straight-Through and BeyondLiyuan Liu, Chengyu Dong, Xiaodong Liu, Bin Yu 等NeurIPS 2023 · 被引用 52 次
- Spectral Smoothing Unveils Phase Transitions in Hierarchical Variational AutoencodersAdeel Pervez, Efstratios GavvesICML 2021 · 被引用 4 次
- Efficient Learning of Discrete-Continuous Computation GraphsDavid Friede, Mathias NiepertNeurIPS 2021 · 被引用 3 次
- Cold Analysis of Rao-Blackwellized Straight-Through Gumbel-Softmax Gradient EstimatorAlexander ShekhovtsovICML 2023 · 被引用 2 次
相关 Paper
- Training Discrete Deep Generative Models via Gapped Straight-Through EstimatorTing-Han Fan, Ta-Chung Chi, Alexander I. Rudnicky, Peter J. RamadgeICML 2022 · 被引用 9 次
- Rao-Blackwellizing the Straight-Through Gumbel-Softmax Gradient EstimatorMax B. Paulus, Chris J. Maddison, Andreas KrauseICLR 2021 · 被引用 48 次
- Path Sample-Analytic Gradient Estimators for Stochastic Binary NetworksAlexander Shekhovtsov, Viktor Yanush, Boris FlachNeurIPS 2020 · 被引用 14 次
- Gradient Estimation with Stochastic Softmax TricksMax B. Paulus, Dami Choi, Daniel Tarlow, Andreas Krause 等NeurIPS 2020 · 被引用 104 次
- Marginalized Stochastic Natural Gradients for Black-Box Variational InferenceGeng Ji, Debora Sujono, Erik B. SudderthICML 2021 · 被引用 9 次
