Effective Estimation of Deep Generative Language Models
Tom Pelsmaeker, Wilker Aziz
Abstract
Advances in variational inference enable parameterisation of probabilistic models by deep neural networks. This combines the statistical transparency of the probabilistic modelling framework with the representational power of deep learning. Yet, due to a problem known as posterior collapse, it is difficult to estimate such models in the context of language modelling effectively. We concentrate on one such model, the variational auto-encoder, which we argue is an important building block in hierarchical probabilistic models of language. This paper contributes a sober view of the problem, a survey of techniques to address it, novel techniques, and extensions to the model. To establish a ranking of techniques, we perform a systematic comparison using Bayesian optimisation and find that many techniques perform reasonably similar, given enough resources. Still, a favourite can be named based on convenience. We also make several empirical observations and recommendations of best practices that should help researchers interested in this exciting field.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext dbbef5c0-5181-46bd-8a1e-5e8806026409Cited by top-tier papers2
- Do sequence-to-sequence VAEs learn global features of sentences?Tom Bosc, Pascal VincentEMNLP 2020 · 5 citations
- Constructing Superior Representations Beyond the Original Documents via a Contrastive Gaussian Fusion Network for ClusteringAo Shen, Ruizhang Huang, Jingjing Xue, Ruina BaiAAAI 2026
Related papers
- Beyond Vanilla Variational Autoencoders: Detecting Posterior Collapse in Conditional and Hierarchical Variational AutoencodersHien Dang, Tho Tran Huu, Tan Minh Nguyen, Nhat HoICLR 2024 · 8 citations
- Posterior Collapse of a Linear Latent Variable ModelZihao Wang, Liu ZiyinNeurIPS 2022 · 29 citations
- A Batch Normalized Inference Network Keeps the KL Vanishing AwayQile Zhu, Wei Bi, Xiaojiang Liu, Xiyao Ma et al.ACL 2020 · 70 citations
- Spectral Smoothing Unveils Phase Transitions in Hierarchical Variational AutoencodersAdeel Pervez, Efstratios GavvesICML 2021 · 4 citations
- Posterior Collapse and Latent Variable Non-identifiabilityYixin Wang, David M. Blei, John P. CunninghamNeurIPS 2021 · 97 citations
