Improving Variational Autoencoders with Density Gap-based Regularization
Jianfei Zhang, Jun Bai, Chenghua Lin, Yanmeng Wang, Wenge Rong
Abstract
Variational autoencoders (VAEs) are one of the most powerful unsupervised learning frameworks in NLP for latent representation learning and latent-directed generation. The classic optimization goal of VAEs is to maximize the Evidence Lower Bound (ELBo), which consists of a conditional likelihood for generation and a negative Kullback-Leibler (KL) divergence for regularization. In practice, optimizing ELBo often leads the posterior distribution of all samples converging to the same degenerated local optimum, namely posterior collapse or KL vanishing. There are effective ways proposed to prevent posterior collapse in VAEs, but we observe that they in essence make trade-offs between posterior collapse and the hole problem, i.e., the mismatch between the aggregated posterior distribution and the prior distribution. To this end, we introduce new training objectives to tackle both problems through a novel regularization based on the probabilistic density gap between the aggregated posterior distribution and the prior distribution. Through experiments on language modeling, latent space visualization, and interpolation, we show that our proposed method can solve both problems effectively and thus outperforms the existing methods in latent-directed generation. To the best of our knowledge, we are the first to jointly solve the hole problem and posterior collapse.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext be1f7dd3-a214-4bbc-98c7-5d6ea8397af5Builds on4
- A Contrastive Learning Approach for Training Variational Autoencoder PriorsJyoti Aneja, Alexander G. Schwing, Jan Kautz, Arash VahdatNeurIPS 2021 · 112 citations
- A Batch Normalized Inference Network Keeps the KL Vanishing AwayQile Zhu, Wei Bi, Xiaojiang Liu, Xiyao Ma et al.ACL 2020 · 70 citations
- Neural Machine Translation with Phrase-Level Universal Visual RepresentationsQingkai Fang, Yang FengACL 2022
- There Are a Thousand Hamlets in a Thousand People's Eyes: Enhancing Knowledge-grounded Dialogue with Personal MemoryTingchen Fu, Xueliang Zhao, Chongyang Tao, Ji-Rong Wen et al.ACL 2022
Related papers
- Posterior Collapse of a Linear Latent Variable ModelZihao Wang, Liu ZiyinNeurIPS 2022 · 29 citations
- Effective Estimation of Deep Generative Language ModelsTom Pelsmaeker, Wilker AzizACL 2020 · 5 citations
- Coupled Variational AutoencoderXiaoran Hao, Patrick ShaftoICML 2023 · 7 citations
- Vector Quantization-Based Regularization for AutoencodersHanwei Wu, Markus FlierlAAAI 2020 · 33 citations
- Generalization Gap in Amortized InferenceMingtian Zhang, Peter Hayes, David BarberNeurIPS 2022 · 14 citations
