On the Uncomputability of Partition Functions in Energy-Based Sequence Models
Chu-Cheng Lin, Arya D. McCarthy
摘要
In this paper, we argue that energy-based sequence models backed by expressive parametric families can result in uncomputable and inapproximable partition functions. Among other things, this makes model selection--and therefore learning model parameters--not only difficult, but generally undecidable. The reason is that there are no good deterministic or randomized estimates of partition functions. Specifically, we exhibit a pathological example where under common assumptions, no useful importance sampling estimates of the partition function can guarantee to have variance bounded below a rational number. As alternatives, we consider sequence model families whose partition functions are computable (if they exist), but at the cost of reduced expressiveness. Our theoretical results suggest that statistical procedures with asymptotic guarantees and sheer (but finite) amounts of compute are not the only things that make sequence modeling work; computability concerns must not be neglected as we consider more expressive model parametrizations.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper2
- HYPRO: A Hybridly Normalized Probabilistic Model for Long-Horizon Prediction of Event SequencesSiqiao Xue, Xiaoming Shi, James Y. Zhang, Hongyuan MeiNeurIPS 2022 · 被引用 65 次
- A Measure-Theoretic Characterization of Tight Language ModelsLi Du, Lucas Torroba Hennigen, Tiago Pimentel, Clara Meister 等ACL 2023 · 被引用 8 次
相关 Paper
- Exposing the Implicit Energy Networks behind Masked Language Models via Metropolis--HastingsKartik Goyal, Chris Dyer, Taylor Berg-KirkpatrickICLR 2022 · 被引用 53 次
- Learning non-Markovian Decision-Making from State-only SequencesAoyang Qin, Feng Gao, Qing Li, Song-Chun Zhu 等NeurIPS 2023 · 被引用 13 次
- Explaining the effects of non-convergent MCMC in the training of Energy-Based ModelsElisabeth Agoritsas, Giovanni Catania, Aurélien Decelle, Beatriz SeoaneICML 2023 · 被引用 17 次
- Efficient Stochastic Optimisation via Sequential Monte CarloJames Cuin, Davide Carbone, Yanbo Tang, O. AkyildizICML 2026 · 被引用 1 次
- SUMO: Unbiased Estimation of Log Marginal Probability for Latent Variable ModelsYucen Luo, Alex Beatson, Mohammad Norouzi, Jun Zhu 等ICLR 2020 · 被引用 29 次
