On the Limitations of Multimodal VAEs
Imant Daunhawer, Thomas M. Sutter, Kieran Chin-Cheong, Emanuele Palumbo, Julia E. Vogt
摘要
Multimodal variational autoencoders (VAEs) have shown promise as efficient generative models for weakly-supervised data. Yet, despite their advantage of weak supervision, they exhibit a gap in generative quality compared to unimodal VAEs, which are completely unsupervised. In an attempt to explain this gap, we uncover a fundamental limitation that applies to a large family of mixture-based multimodal VAEs. We prove that the sub-sampling of modalities enforces an undesirable upper bound on the multimodal ELBO and thereby limits the generative quality of the respective models. Empirically, we showcase the generative quality gap on both synthetic and real data and present the tradeoffs between different variants of multimodal VAEs. We find that none of the existing approaches fulfills all desired criteria of an effective multimodal generative model when applied on more complex datasets than those used in previous benchmarks. In summary, we identify, formalize, and validate fundamental limitations of VAE-based approaches for modeling weakly-supervised data and discuss implications for real-world applications.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- HALC: Object Hallucination Reduction via Adaptive Focal-Contrast DecodingZhaorun Chen, Zhuokai Zhao, Hongyin Luo, Huaxiu Yao 等ICML 2024 · 被引用 164 次
- Deep Generative Clustering with Multimodal Diffusion Variational AutoencodersEmanuele Palumbo, Laura Manduchi, Sonia Laguna, Daphné Chopard 等ICLR 2024 · 被引用 21 次
- Disentanglement of Variations with Multimodal Generative ModelingYijie Zhang, Yiyang Shen, Weiran WangICLR 2026 · 被引用 6 次
- Disentangled Cross-Modal Representation Learning with Enhanced Mutual SupervisionLu Gao, Wenlan Chen, Daoyuan Wang, Fei Guo 等NeurIPS 2025 · 被引用 5 次
- CLAP: Collaborative Adaptation for Patchwork LearningSen Cui, Abudukelimu Wuerkaixi, Weishen Pan, Jian Liang 等ICLR 2024 · 被引用 3 次
它引用的顶会 Paper8
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray 等ICML 2021 · 被引用 6,356 次
- Few-Shot Unsupervised Image-to-Image TranslationMing-Yu Liu, Xun Huang, Arun Mallya, Tero Karras 等ICCV 2019 · 被引用 668 次
- Self-Supervised Learning with Data Augmentations Provably Isolates Content from StyleJulius von Kügelgen, Yash Sharma, Luigi Gresele, Wieland Brendel 等NeurIPS 2021 · 被引用 421 次
- Weakly-Supervised Disentanglement Without CompromisesFrancesco Locatello, Ben Poole, Gunnar Rätsch, Bernhard Schölkopf 等ICML 2020 · 被引用 361 次
- Generalized Multimodal ELBOThomas M. Sutter, Imant Daunhawer, Julia E. VogtICLR 2021 · 被引用 130 次
相关 Paper
- MMVAE+: Enhancing the Generative Quality of Multimodal VAEs without CompromisesEmanuele Palumbo, Imant Daunhawer, Julia E. VogtICLR 2023
- Aggregation of Dependent Expert Distributions in Multimodal Variational AutoencodersRogelio Andrade Mancisidor, Robert Jenssen, Shujian Yu, Michael KampffmeyerICML 2025
- Multimodal Gaussian Mixture Variational Autoencoder with Consistency RegularizationsYarui Chen, Lehan Hong, Jianlin Shao, Jianning Yang 等AAAI 2026
- Efficient Modality Translation via Arbitrary Conditioning and Wasserstein RegularizationTomás Tokár, Scott SannerAAAI 2026
- Unity by Diversity: Improved Representation Learning for Multimodal VAEsThomas M. Sutter, Yang Meng, Andrea Agostini, Daphné Chopard 等NeurIPS 2024 · 被引用 21 次
