A Batch Normalized Inference Network Keeps the KL Vanishing Away
Qile Zhu, Wei Bi, Xiaojiang Liu, Xiyao Ma, Xiaolin Li, Dapeng Wu
Abstract
Variational Autoencoder (VAE) is widely used as a generative model to approximate a model's posterior on latent variables by combining the amortized variational inference and deep neural networks. However, when paired with strong autoregressive decoders, VAE often converges to a degenerated local optimum known as "posterior collapse". Previous approaches consider the Kullback-Leibler divergence (KL) individual for each datapoint. We propose to let the KL follow a distribution across the whole dataset, and analyze that it is sufficient to prevent posterior collapse by keeping the expectation of the KL's distribution positive. Then we propose Batch Normalized-VAE (BN-VAE), a simple but effective approach to set a lower bound of the expectation by regularizing the distribution of the approximate posterior's parameters. Without introducing any new model component or modifying the objective, our approach can avoid the posterior collapse effectively and efficiently. We further show that the proposed BN-VAE can be extended to conditional VAE (CVAE). Empirically, our approach surpasses strong autoregressive baselines on language modeling, text classification and dialogue generation, and rivals more complex approaches while keeping almost the same training time as VAE.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2d354df5-c26d-4afd-a6f4-554a7d64cd73Cited by top-tier papers7
- Optimus: Organizing Sentences via Pre-trained Modeling of a Latent SpaceChunyuan Li, Xiang Gao, Yuan Li, Baolin Peng et al.EMNLP 2020 · 132 citations
- Conditional Variational Autoencoder for Sign Language Translation with Cross-Modal AlignmentRui Zhao, Liang Zhang, Biao Fu, Cong Hu et al.AAAI 2024 · 36 citations
- Learning Agent Representations for Ice HockeyGuiliang Liu, Oliver Schulte, Pascal Poupart, Mike Rudd et al.NeurIPS 2020 · 13 citations
- Improving Variational Autoencoders with Density Gap-based RegularizationJianfei Zhang, Jun Bai, Chenghua Lin, Yanmeng Wang et al.NeurIPS 2022 · 11 citations
- Towards Diverse, Relevant and Coherent Open-Domain Dialogue Generation via Hybrid Latent VariablesBin Sun, Yitong Li, Fei Mi, Weichao Wang et al.AAAI 2023 · 8 citations
Related papers
- Posterior Collapse of a Linear Latent Variable ModelZihao Wang, Liu ZiyinNeurIPS 2022 · 29 citations
- Effective Estimation of Deep Generative Language ModelsTom Pelsmaeker, Wilker AzizACL 2020 · 5 citations
- On the Encoder-Decoder Incompatibility in Variational Text Modeling and BeyondChen Wu, Prince Zizhuang Wang, William Yang WangACL 2020 · 2 citations
- Beyond Vanilla Variational Autoencoders: Detecting Posterior Collapse in Conditional and Hierarchical Variational AutoencodersHien Dang, Tho Tran Huu, Tan Minh Nguyen, Nhat HoICLR 2024 · 8 citations
- Vector Quantization-Based Regularization for AutoencodersHanwei Wu, Markus FlierlAAAI 2020 · 33 citations
