Generalized Multimodal ELBO
Thomas M. Sutter, Imant Daunhawer, Julia E. Vogt
Abstract
Multiple data types naturally co-occur when describing real-world phenomena and learning from them is a long-standing goal in machine learning research. However, existing self-supervised generative models approximating an ELBO are not able to fulfill all desired requirements of multimodal models: their posterior approximation functions lead to a trade-off between the semantic coherence and the ability to learn the joint data distribution. We propose a new, generalized ELBO formulation for multimodal data that overcomes these limitations. The new objective encompasses two previous methods as special cases and combines their benefits without compromises. In extensive experiments, we demonstrate the advantage of the proposed method compared to state-of-the-art models in self-supervised, generative learning tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fbf233f0-8dad-4b69-bef6-ee8c13e092d1Cited by top-tier papers34
- Visual Decoding and Reconstruction via EEG Embeddings with Guided DiffusionDongyang Li, Chen Wei, Shiying Li, Jiachen Zou et al.NeurIPS 2024 · 164 citations
- On the Limitations of Multimodal VAEsImant Daunhawer, Thomas M. Sutter, Kieran Chin-Cheong, Emanuele Palumbo et al.ICLR 2022 · 50 citations
- Mitigating Modality Collapse in Multimodal VAEs via Impartial OptimizationAdrián Javaloy, Maryam Meghdadi, Isabel ValeraICML 2022 · 49 citations
- Deep Variational Incomplete Multi-View Clustering: Exploring Shared Clustering StructuresGehui Xu, Jie Wen, Chengliang Liu, Bing Hu et al.AAAI 2024 · 44 citations
- Learning Multimodal VAEs through Mutual SupervisionTom Joy, Yuge Shi, Philip H. S. Torr, Tom Rainforth et al.ICLR 2022 · 27 citations
Builds on2
- Multimodal Generative Learning Utilizing Jensen-Shannon-DivergenceThomas M. Sutter, Imant Daunhawer, Julia E. VogtNeurIPS 2020 · 105 citations
- Relating by Contrasting: A Data-efficient Framework for Multimodal Generative ModelsYuge Shi, Brooks Paige, Philip H. S. Torr, N. SiddharthICLR 2021 · 42 citations
Related papers
- Efficient Modality Translation via Arbitrary Conditioning and Wasserstein RegularizationTomás Tokár, Scott SannerAAAI 2026
- Unity by Diversity: Improved Representation Learning for Multimodal VAEsThomas M. Sutter, Yang Meng, Andrea Agostini, Daphné Chopard et al.NeurIPS 2024 · 21 citations
- Multimodal Adversarially Learned Inference with Factorized DiscriminatorsWenxue Chen, Jianke ZhuAAAI 2022 · 3 citations
- Conditional Generative Modeling via Learning the Latent SpaceSameera Ramasinghe, Kanchana Nisal Ranasinghe, Salman H. Khan, Nick Barnes et al.ICLR 2021 · 10 citations
- Multimodal Variational Autoencoder: A Barycentric ViewPeijie Qiu, Wenhui Zhu, Sayantan Kumar, Xiwen Chen et al.AAAI 2025 · 2 citations
