A Batch Normalized Inference Network Keeps the KL Vanishing Away
Qile Zhu, Wei Bi, Xiaojiang Liu, Xiyao Ma, Xiaolin Li, Dapeng Wu
摘要
Variational Autoencoder (VAE) is widely used as a generative model to approximate a model's posterior on latent variables by combining the amortized variational inference and deep neural networks. However, when paired with strong autoregressive decoders, VAE often converges to a degenerated local optimum known as "posterior collapse". Previous approaches consider the Kullback-Leibler divergence (KL) individual for each datapoint. We propose to let the KL follow a distribution across the whole dataset, and analyze that it is sufficient to prevent posterior collapse by keeping the expectation of the KL's distribution positive. Then we propose Batch Normalized-VAE (BN-VAE), a simple but effective approach to set a lower bound of the expectation by regularizing the distribution of the approximate posterior's parameters. Without introducing any new model component or modifying the objective, our approach can avoid the posterior collapse effectively and efficiently. We further show that the proposed BN-VAE can be extended to conditional VAE (CVAE). Empirically, our approach surpasses strong autoregressive baselines on language modeling, text classification and dialogue generation, and rivals more complex approaches while keeping almost the same training time as VAE.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Optimus: Organizing Sentences via Pre-trained Modeling of a Latent SpaceChunyuan Li, Xiang Gao, Yuan Li, Baolin Peng 等EMNLP 2020 · 被引用 132 次
- Conditional Variational Autoencoder for Sign Language Translation with Cross-Modal AlignmentRui Zhao, Liang Zhang, Biao Fu, Cong Hu 等AAAI 2024 · 被引用 36 次
- Learning Agent Representations for Ice HockeyGuiliang Liu, Oliver Schulte, Pascal Poupart, Mike Rudd 等NeurIPS 2020 · 被引用 13 次
- Improving Variational Autoencoders with Density Gap-based RegularizationJianfei Zhang, Jun Bai, Chenghua Lin, Yanmeng Wang 等NeurIPS 2022 · 被引用 11 次
- Towards Diverse, Relevant and Coherent Open-Domain Dialogue Generation via Hybrid Latent VariablesBin Sun, Yitong Li, Fei Mi, Weichao Wang 等AAAI 2023 · 被引用 8 次
相关 Paper
- Posterior Collapse of a Linear Latent Variable ModelZihao Wang, Liu ZiyinNeurIPS 2022 · 被引用 29 次
- Effective Estimation of Deep Generative Language ModelsTom Pelsmaeker, Wilker AzizACL 2020 · 被引用 5 次
- On the Encoder-Decoder Incompatibility in Variational Text Modeling and BeyondChen Wu, Prince Zizhuang Wang, William Yang WangACL 2020 · 被引用 2 次
- Beyond Vanilla Variational Autoencoders: Detecting Posterior Collapse in Conditional and Hierarchical Variational AutoencodersHien Dang, Tho Tran Huu, Tan Minh Nguyen, Nhat HoICLR 2024 · 被引用 8 次
- Vector Quantization-Based Regularization for AutoencodersHanwei Wu, Markus FlierlAAAI 2020 · 被引用 33 次
