On the Uncomputability of Partition Functions in Energy-Based Sequence Models
Chu-Cheng Lin, Arya D. McCarthy
Abstract
In this paper, we argue that energy-based sequence models backed by expressive parametric families can result in uncomputable and inapproximable partition functions. Among other things, this makes model selection--and therefore learning model parameters--not only difficult, but generally undecidable. The reason is that there are no good deterministic or randomized estimates of partition functions. Specifically, we exhibit a pathological example where under common assumptions, no useful importance sampling estimates of the partition function can guarantee to have variance bounded below a rational number. As alternatives, we consider sequence model families whose partition functions are computable (if they exist), but at the cost of reduced expressiveness. Our theoretical results suggest that statistical procedures with asymptotic guarantees and sheer (but finite) amounts of compute are not the only things that make sequence modeling work; computability concerns must not be neglected as we consider more expressive model parametrizations.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get bf1cb2d8-c37a-47b1-843b-88096348e779Cited by top-tier papers2
- HYPRO: A Hybridly Normalized Probabilistic Model for Long-Horizon Prediction of Event SequencesSiqiao Xue, Xiaoming Shi, James Y. Zhang, Hongyuan MeiNeurIPS 2022 · 65 citations
- A Measure-Theoretic Characterization of Tight Language ModelsLi Du, Lucas Torroba Hennigen, Tiago Pimentel, Clara Meister et al.ACL 2023 · 8 citations
Related papers
- Exposing the Implicit Energy Networks behind Masked Language Models via Metropolis--HastingsKartik Goyal, Chris Dyer, Taylor Berg-KirkpatrickICLR 2022 · 53 citations
- Learning non-Markovian Decision-Making from State-only SequencesAoyang Qin, Feng Gao, Qing Li, Song-Chun Zhu et al.NeurIPS 2023 · 13 citations
- Explaining the effects of non-convergent MCMC in the training of Energy-Based ModelsElisabeth Agoritsas, Giovanni Catania, Aurélien Decelle, Beatriz SeoaneICML 2023 · 17 citations
- Efficient Stochastic Optimisation via Sequential Monte CarloJames Cuin, Davide Carbone, Yanbo Tang, O. AkyildizICML 2026 · 1 citation
- SUMO: Unbiased Estimation of Log Marginal Probability for Latent Variable ModelsYucen Luo, Alex Beatson, Mohammad Norouzi, Jun Zhu et al.ICLR 2020 · 29 citations
