MMVAE+: Enhancing the Generative Quality of Multimodal VAEs without Compromises
Emanuele Palumbo, Imant Daunhawer, Julia E. Vogt
摘要
Multimodal VAEs have recently gained attention as efficient models for weaklysupervised generative learning with a large number of modalities. However, all existing variants of multimodal VAEs are affected by a non-trivial trade-off between generative quality and generative coherence. We focus on the mixture-ofexperts multimodal VAE (MMVAE), which achieves good coherence only at the expense of sample diversity and a resulting lack of generative quality. We present a novel variant of the MMVAE that improves its generative quality, while maintaining high semantic coherence. For this, shared and modality-specific information is modelled in separate latent subspaces. In contrast to previous approaches with separate subspaces, our model is robust to changes in latent dimensionality and regularization hyperparameters. We show that our model achieves both good generative coherence and high generative quality in challenging experiments, including more complex multimodal datasets than those used in previous works.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- Deep Generative Clustering with Multimodal Diffusion Variational AutoencodersEmanuele Palumbo, Laura Manduchi, Sonia Laguna, Daphné Chopard 等ICLR 2024 · 被引用 21 次
- Unity by Diversity: Improved Representation Learning for Multimodal VAEsThomas M. Sutter, Yang Meng, Andrea Agostini, Daphné Chopard 等NeurIPS 2024 · 被引用 21 次
- Disentanglement of Variations with Multimodal Generative ModelingYijie Zhang, Yiyang Shen, Weiran WangICLR 2026 · 被引用 6 次
- Multi-Modal Latent Variables for Cross-Individual Primary Visual Cortex Modeling and AnalysisYu Zhu, Bo Lei, Chunfeng Song, Wanli Ouyang 等AAAI 2025 · 被引用 5 次
- Disentangled Cross-Modal Representation Learning with Enhanced Mutual SupervisionLu Gao, Wenlan Chen, Daoyuan Wang, Fei Guo 等NeurIPS 2025 · 被引用 5 次
它引用的顶会 Paper2
相关 Paper
- Hölder++: Improving Quality-Coherence Trade-off in Multimodal VAEsHuyen Vo, María Martínez-García, Isabel ValeraICML 2026
- Efficient Modality Translation via Arbitrary Conditioning and Wasserstein RegularizationTomás Tokár, Scott SannerAAAI 2026
- Learning Multimodal VAEs through Mutual SupervisionTom Joy, Yuge Shi, Philip H. S. Torr, Tom Rainforth 等ICLR 2022 · 被引用 27 次
- ShaLa: Multimodal Shared Latent Generative ModellingJiali Cui, Yan-Ying Chen, Yanxia Zhang, Matthew KlenkAAAI 2026
- Multimodal Gaussian Mixture Variational Autoencoder with Consistency RegularizationsYarui Chen, Lehan Hong, Jianlin Shao, Jianning Yang 等AAAI 2026
