MMVAE+: Enhancing the Generative Quality of Multimodal VAEs without Compromises
Emanuele Palumbo, Imant Daunhawer, Julia E. Vogt
Abstract
Multimodal VAEs have recently gained attention as efficient models for weaklysupervised generative learning with a large number of modalities. However, all existing variants of multimodal VAEs are affected by a non-trivial trade-off between generative quality and generative coherence. We focus on the mixture-ofexperts multimodal VAE (MMVAE), which achieves good coherence only at the expense of sample diversity and a resulting lack of generative quality. We present a novel variant of the MMVAE that improves its generative quality, while maintaining high semantic coherence. For this, shared and modality-specific information is modelled in separate latent subspaces. In contrast to previous approaches with separate subspaces, our model is robust to changes in latent dimensionality and regularization hyperparameters. We show that our model achieves both good generative coherence and high generative quality in challenging experiments, including more complex multimodal datasets than those used in previous works.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c1c646a7-cd00-4457-ae3b-d3b9c94a9dd3Cited by top-tier papers14
- Deep Generative Clustering with Multimodal Diffusion Variational AutoencodersEmanuele Palumbo, Laura Manduchi, Sonia Laguna, Daphné Chopard et al.ICLR 2024 · 21 citations
- Unity by Diversity: Improved Representation Learning for Multimodal VAEsThomas M. Sutter, Yang Meng, Andrea Agostini, Daphné Chopard et al.NeurIPS 2024 · 21 citations
- Disentanglement of Variations with Multimodal Generative ModelingYijie Zhang, Yiyang Shen, Weiran WangICLR 2026 · 6 citations
- Multi-Modal Latent Variables for Cross-Individual Primary Visual Cortex Modeling and AnalysisYu Zhu, Bo Lei, Chunfeng Song, Wanli Ouyang et al.AAAI 2025 · 5 citations
- Disentangled Cross-Modal Representation Learning with Enhanced Mutual SupervisionLu Gao, Wenlan Chen, Daoyuan Wang, Fei Guo et al.NeurIPS 2025 · 5 citations
Builds on2
Related papers
- Hölder++: Improving Quality-Coherence Trade-off in Multimodal VAEsHuyen Vo, María Martínez-García, Isabel ValeraICML 2026
- Efficient Modality Translation via Arbitrary Conditioning and Wasserstein RegularizationTomás Tokár, Scott SannerAAAI 2026
- Learning Multimodal VAEs through Mutual SupervisionTom Joy, Yuge Shi, Philip H. S. Torr, Tom Rainforth et al.ICLR 2022 · 27 citations
- ShaLa: Multimodal Shared Latent Generative ModellingJiali Cui, Yan-Ying Chen, Yanxia Zhang, Matthew KlenkAAAI 2026
- Multimodal Gaussian Mixture Variational Autoencoder with Consistency RegularizationsYarui Chen, Lehan Hong, Jianlin Shao, Jianning Yang et al.AAAI 2026
