Multimodal Variational Autoencoder: A Barycentric View
Peijie Qiu, Wenhui Zhu, Sayantan Kumar, Xiwen Chen, Jin Yang, Xiaotong Sun, Abolfazl Razi, Yalin Wang, Aristeidis Sotiras
Abstract
Multiple signal modalities, such as vision and sounds, are naturally present in real-world phenomena. Recently, there has been growing interest in learning generative models, in particular variational autoencoder (VAE), for multimodal representation learning especially in the case of missing modalities. The primary goal of these models is to learn a modality-invariant and modality-specific representation that characterizes information across multiple modalities. Previous attempts at multimodal VAEs approach this mainly through the lens of experts, aggregating unimodal inference distributions with a product of experts (PoE), a mixture of experts (MoE), or a combination of both. In this paper, we provide an alternative generic and theoretical formulation of multimodal VAE through the lens of barycenter. We first show that PoE and MoE are specific instances of barycenters, derived by minimizing the asymmetric weighted KL divergence to unimodal inference distributions. Our novel formulation extends these two barycenters to a more flexible choice by considering different types of divergences. In particular, we explore the Wasserstein barycenter defined by the 2-Wasserstein distance, which better preserves the geometry of unimodal distributions by capturing both modality-specific and modality-invariant representations compared to KL divergence. Empirical studies on three multimodal benchmarks demonstrated the effectiveness of the proposed method.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- Geometry-Aware Variational Information Maximization for Deep Incomplete Multi-view ClusteringWenlan Chen, Lu Gao, Daoyuan Wang, Fei Guo et al.AAAI 2026
- Hölder++: Improving Quality-Coherence Trade-off in Multimodal VAEsHuyen Vo, María Martínez-García, Isabel ValeraICML 2026
- Multimodal Gaussian Mixture Variational Autoencoder with Consistency RegularizationsYarui Chen, Lehan Hong, Jianlin Shao, Jianning Yang et al.AAAI 2026
- Information-Theoretic Disentangled Latent Modeling with Conditional Diffusion for Incomplete Multi-View ClusteringWenlan Chen, Lu Gao, Daoyuan Wang, Cheng Liang et al.ICML 2026
- Gated Variational Graph Autoencoders as Experts with Competition and Consensus for Multi-view ClusteringZhaoliang Chen, William K. Cheung, Hong-Ning Dai, Byron Choi et al.AAAI 2026
Builds on3
- Generalized Multimodal ELBOThomas M. Sutter, Imant Daunhawer, Julia E. VogtICLR 2021 · 130 citations
- MMVAE+: Enhancing the Generative Quality of Multimodal VAEs without CompromisesEmanuele Palumbo, Imant Daunhawer, Julia E. VogtICLR 2023
- Vision Transformers are Parameter-Efficient Audio-Visual LearnersYan-Bo Lin, Yi-Lin Sung, Jie Lei, Mohit Bansal et al.CVPR 2023
Related papers
- Efficient Modality Translation via Arbitrary Conditioning and Wasserstein RegularizationTomás Tokár, Scott SannerAAAI 2026
- Unity by Diversity: Improved Representation Learning for Multimodal VAEsThomas M. Sutter, Yang Meng, Andrea Agostini, Daphné Chopard et al.NeurIPS 2024 · 21 citations
- Aggregation of Dependent Expert Distributions in Multimodal Variational AutoencodersRogelio Andrade Mancisidor, Robert Jenssen, Shujian Yu, Michael KampffmeyerICML 2025
- Deep Variational Incomplete Multi-View Clustering with Information-Theoretic GuidanceWenlan Chen, Lu Gao, Cheng Liang, Fei GuoACM MM 2025 · 1 citation
- Associative Variational Auto-Encoder with Distributed Latent Spaces and AssociatorsDae Ung Jo, Byeongju Lee, Jongwon Choi, Haanju Yoo et al.AAAI 2020 · 8 citations
