A VAE for Transformers with Nonparametric Variational Information Bottleneck
James Henderson, Fabio Fehr
Abstract
We propose a Variational AutoEncoder (VAE) for Transformers by developing a Variational Information Bottleneck (VIB) regulariser for Transformer embeddings. We formalise such attention-based representations as mixture distributions, and use Bayesian nonparametrics to develop a Nonparametric VIB (NVIB) for them. The variable number of mixture components supported by nonparametrics captures the variable number of vectors supported by attention, and exchangeable distributions from nonparametrics capture the permutation invariance of attention. Our Transformer VAE (NVAE) uses NVIB to regularise the information passing from the Transformer encoder to the Transformer decoder. Evaluations of a NVAE, trained on natural language text, demonstrate that NVIB can regularise the number of mixture components in the induced embedding whilst maintaining generation quality and reconstruction capacity.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers3
- NOIR: Privacy-Preserving Generation of Code with Open-Source LLMsKhoa Nguyen, Khiem Ton, NhatHai Phan, Issa Khalil et al.USENIX Security 2026 · 2 citations
- Measuring In-Context Computation Complexity via Hidden State PredictionVincent Herrmann, Róbert Csordás, Jürgen SchmidhuberICML 2025
- Compositional Generalization through Gradient Search in Nonparametric Latent SpaceHaruki Shirakami, James HendersonICLR 2026
Related papers
- Improving Transformers with Probabilistic Attention KeysTam Minh Nguyen, Tan Minh Nguyen, Dung D. Le, Duy Khuong Nguyen et al.ICML 2022 · 38 citations
- Generalization Guarantees for Representation Learning via Data-Dependent Gaussian Mixture PriorsMilad Sefidgaran, Abdellatif Zaidi, Piotr KrasnowskiICLR 2025
- Vector Quantization-Based Regularization for AutoencodersHanwei Wu, Markus FlierlAAAI 2020 · 33 citations
- Evade the Trap of Mediocrity: Promoting Diversity and Novelty in Text Generation via Concentrating AttentionWenhao Li, Xiaoyuan Yi, Jinyi Hu, Maosong Sun et al.EMNLP 2022 · 1 citation
- BooVAE: Boosting Approach for Continual Learning of VAEEvgenii Egorov, Anna Kuzina, Evgeny BurnaevNeurIPS 2021 · 34 citations
