Generalized Multimodal ELBO
Thomas M. Sutter, Imant Daunhawer, Julia E. Vogt
摘要
Multiple data types naturally co-occur when describing real-world phenomena and learning from them is a long-standing goal in machine learning research. However, existing self-supervised generative models approximating an ELBO are not able to fulfill all desired requirements of multimodal models: their posterior approximation functions lead to a trade-off between the semantic coherence and the ability to learn the joint data distribution. We propose a new, generalized ELBO formulation for multimodal data that overcomes these limitations. The new objective encompasses two previous methods as special cases and combines their benefits without compromises. In extensive experiments, we demonstrate the advantage of the proposed method compared to state-of-the-art models in self-supervised, generative learning tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper34
- Visual Decoding and Reconstruction via EEG Embeddings with Guided DiffusionDongyang Li, Chen Wei, Shiying Li, Jiachen Zou 等NeurIPS 2024 · 被引用 164 次
- On the Limitations of Multimodal VAEsImant Daunhawer, Thomas M. Sutter, Kieran Chin-Cheong, Emanuele Palumbo 等ICLR 2022 · 被引用 50 次
- Mitigating Modality Collapse in Multimodal VAEs via Impartial OptimizationAdrián Javaloy, Maryam Meghdadi, Isabel ValeraICML 2022 · 被引用 49 次
- Deep Variational Incomplete Multi-View Clustering: Exploring Shared Clustering StructuresGehui Xu, Jie Wen, Chengliang Liu, Bing Hu 等AAAI 2024 · 被引用 44 次
- Learning Multimodal VAEs through Mutual SupervisionTom Joy, Yuge Shi, Philip H. S. Torr, Tom Rainforth 等ICLR 2022 · 被引用 27 次
它引用的顶会 Paper2
相关 Paper
- Efficient Modality Translation via Arbitrary Conditioning and Wasserstein RegularizationTomás Tokár, Scott SannerAAAI 2026
- Unity by Diversity: Improved Representation Learning for Multimodal VAEsThomas M. Sutter, Yang Meng, Andrea Agostini, Daphné Chopard 等NeurIPS 2024 · 被引用 21 次
- Multimodal Adversarially Learned Inference with Factorized DiscriminatorsWenxue Chen, Jianke ZhuAAAI 2022 · 被引用 3 次
- Conditional Generative Modeling via Learning the Latent SpaceSameera Ramasinghe, Kanchana Nisal Ranasinghe, Salman H. Khan, Nick Barnes 等ICLR 2021 · 被引用 10 次
- Multimodal Variational Autoencoder: A Barycentric ViewPeijie Qiu, Wenhui Zhu, Sayantan Kumar, Xiwen Chen 等AAAI 2025 · 被引用 2 次
