Scaling Structured Inference with Randomization
Yao Fu, John P. Cunningham, Mirella Lapata
摘要
Deep discrete structured models have seen considerable progress recently, but traditional inference using dynamic programming (DP) typically works with a small number of states (less than hundreds), which severely limits model capacity. At the same time, across machine learning, there is a recent trend of using randomized truncation techniques to accelerate computations involving large sums. Here, we propose a family of randomized dynamic programming (RDP) algorithms for scaling structured models to tens of thousands of latent states. Our method is widely applicable to classical DP-based inference (partition, marginal, reparameterization, entropy) and different graph structures (chains, trees, and more general hypergraphs). It is also compatible with automatic differentiation: it can be integrated with neural networks seamlessly and learned with gradient-based optimizers. Our core technique approximates the sum-product by restricting and reweighting DP on a small subset of nodes, which reduces computation by orders of magnitude. We further achieve low bias and variance via Rao-Blackwellization and importance sampling. Experiments over different graphs demonstrate the accuracy and efficiency of our approach. Furthermore, when using RDP for training a structured variational autoencoder with a scaled inference network, we achieve better test likelihood than baselines and successfully prevent posterior collapse. code at: https://github.com/FranxYao/RDP
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper7
- Gradient Estimation with Stochastic Softmax TricksMax B. Paulus, Dami Choi, Daniel Tarlow, Andreas Krause 等NeurIPS 2020 · 被引用 104 次
- Efficient Second-Order TreeCRF for Neural Dependency ParsingYu Zhang, Zhenghua Li, Min ZhangACL 2020 · 被引用 90 次
- Randomized Automatic DifferentiationDeniz Oktay, Nick McGreivy, Joshua Aduol, Alex Beatson 等ICLR 2021 · 被引用 31 次
- Efficient Marginalization of Discrete and Structured Latent Variables via SparsityGonçalo M. Correia, Vlad Niculae, Wilker Aziz, André F. T. MartinsNeurIPS 2020 · 被引用 25 次
- Bias-Free Scalable Gaussian Processes via Randomized TruncationsAndres Potapczynski, Luhuan Wu, Dan Biderman, Geoff Pleiss 等ICML 2021 · 被引用 23 次
相关 Paper
- Amortized Population Gibbs Samplers with Neural Sufficient StatisticsHao Wu, Heiko Zimmermann, Eli Sennesh, Tuan Anh Le 等ICML 2020 · 被引用 7 次
- Effective Estimation of Deep Generative Language ModelsTom Pelsmaeker, Wilker AzizACL 2020 · 被引用 5 次
- Revisiting Structured Variational AutoencodersYixiu Zhao, Scott W. LindermanICML 2023 · 被引用 15 次
- Top-Down Bayesian Posterior Sampling for Sum-Product NetworksSoma Yokoi, Issei SatoKDD 2024
- Probabilistic Circuits for Variational Inference in Discrete Graphical ModelsAndy Shih, Stefano ErmonNeurIPS 2020 · 被引用 29 次
