Leveraging Recursive Gumbel-Max Trick for Approximate Inference in Combinatorial Spaces
Kirill Struminsky, Artyom Gadetsky, Denis Rakitin, Danil Karpushkin, Dmitry P. Vetrov
摘要
Structured latent variables allow incorporating meaningful prior knowledge into deep learning models. However, learning with such variables remains challenging because of their discrete nature. Nowadays, the standard learning approach is to define a latent variable as a perturbed algorithm output and to use a differentiable surrogate for training. In general, the surrogate puts additional constraints on the model and inevitably leads to biased gradients. To alleviate these shortcomings, we extend the Gumbel-Max trick to define distributions over structured domains. We avoid the differentiable surrogates by leveraging the score function estimators for optimization. In particular, we highlight a family of recursive algorithms with a common feature we call stochastic invariant. The feature allows us to construct reliable gradient estimates and control variates without additional constraints on the model. In our experiments, we consider various structured latent variable models and achieve results competitive with relaxation-based counterparts.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Unsupervised Learning for Combinatorial Optimization with Principled Objective RelaxationHaoyu Wang, Nan Wu, Hang Yang, Cong Hao 等NeurIPS 2022 · 被引用 54 次
- GeoPhy: Differentiable Phylogenetic Inference via Geometric Gradients of Tree TopologiesTakahiro Mimori, Michiaki HamadaNeurIPS 2023 · 被引用 17 次
- Differentiable Clustering with Perturbed Spanning ForestsLawrence Stewart, Francis R. Bach, Felipe Llinares-López, Quentin BerthetNeurIPS 2023 · 被引用 16 次
- Noise-Resilient Symbolic Regression with Dynamic Gating Reinforcement LearningChenglu Sun, Shuo Shen, Wenzhi Tao, Deyi Xue 等AAAI 2025 · 被引用 5 次
- Latent Optimal Paths by Gumbel Propagation for Variational Bayesian Dynamic ProgrammingXinlei Niu, Christian Walder, Jing Zhang, Charles Patrick MartinICML 2024
它引用的顶会 Paper8
- Gradient Estimation with Stochastic Softmax TricksMax B. Paulus, Dami Choi, Daniel Tarlow, Andreas Krause 等NeurIPS 2020 · 被引用 104 次
- VarGrad: A Low-Variance Gradient Estimator for Variational InferenceLorenz Richter, Ayman Boustati, Nikolas Nüsken, Francisco J. R. Ruiz 等NeurIPS 2020 · 被引用 90 次
- Rao-Blackwellizing the Straight-Through Gumbel-Softmax Gradient EstimatorMax B. Paulus, Chris J. Maddison, Andreas KrauseICLR 2021 · 被引用 48 次
- Discovering Non-monotonic Autoregressive Orderings with Variational InferenceXuanlin Li, Brandon Trabucco, Dong Huk Park, Michael Luo 等ICLR 2021 · 被引用 17 次
- Latent Template Induction with Gumbel-CRFsYao Fu, Chuanqi Tan, Bin Bi, Mosha Chen 等NeurIPS 2020 · 被引用 15 次
相关 Paper
- Efficient Marginalization of Discrete and Structured Latent Variables via SparsityGonçalo M. Correia, Vlad Niculae, Wilker Aziz, André F. T. MartinsNeurIPS 2020 · 被引用 25 次
- Low-Variance Black-Box Gradient Estimates for the Plackett-Luce DistributionArtyom Gadetsky, Kirill Struminsky, Christopher Robinson, Novi Quadrianto 等AAAI 2020 · 被引用 11 次
- Learning Permutation from Structure Without SupervisionRan Eisenberg, Ofir LindenbaumICML 2026
- Learning Generalized Gumbel-max Causal MechanismsGuy Lorberbom, Daniel D. Johnson, Chris J. Maddison, Daniel Tarlow 等NeurIPS 2021 · 被引用 25 次
- Cold Analysis of Rao-Blackwellized Straight-Through Gumbel-Softmax Gradient EstimatorAlexander ShekhovtsovICML 2023 · 被引用 2 次
