Learning Multimodal VAEs through Mutual Supervision
Tom Joy, Yuge Shi, Philip H. S. Torr, Tom Rainforth, Sebastian M. Schmon, Siddharth Narayanaswamy
Abstract
Multimodal variational autoencoders (VAEs) seek to model the joint distribution over heterogeneous data (e.g. vision, language), whilst also capturing a shared representation across such modalities. Prior work has typically combined information from the modalities by reconciling idiosyncratic representations directly in the recognition model through explicit products, mixtures, or other such factorisations. Here we introduce a novel alternative, the Mutually supErvised Multimodal VAE (MEME), that avoids such explicit combinations by repurposing semisupervised VAEs to combine information between modalities implicitly through mutual supervision. This formulation naturally allows learning from partiallyobserved data where some modalities can be entirely missing-something that most existing approaches either cannot handle, or do so to a limited extent. We demonstrate that MEME outperforms baselines on standard metrics across both partial and complete observation schemes on the MNIST-SVHN (image-image) and CUB (image-text) datasets 1 . We also contrast the quality of the representations learnt by mutual supervision against standard approaches and observe interesting trends in its ability to capture relatedness between data.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers9
- Multimodal Variational Auto-encoder based Audio-Visual SegmentationYuxin Mao, Jing Zhang, Mochu Xiang, Yiran Zhong et al.ICCV 2023 · 57 citations
- Deep Generative Clustering with Multimodal Diffusion Variational AutoencodersEmanuele Palumbo, Laura Manduchi, Sonia Laguna, Daphné Chopard et al.ICLR 2024 · 21 citations
- Unity by Diversity: Improved Representation Learning for Multimodal VAEsThomas M. Sutter, Yang Meng, Andrea Agostini, Daphné Chopard et al.NeurIPS 2024 · 21 citations
- Multi-Modal Latent Variables for Cross-Individual Primary Visual Cortex Modeling and AnalysisYu Zhu, Bo Lei, Chunfeng Song, Wanli Ouyang et al.AAAI 2025 · 5 citations
- Disentangled Cross-Modal Representation Learning with Enhanced Mutual SupervisionLu Gao, Wenlan Chen, Daoyuan Wang, Fei Guo et al.NeurIPS 2025 · 5 citations
Builds on4
- Generalized Multimodal ELBOThomas M. Sutter, Imant Daunhawer, Julia E. VogtICLR 2021 · 130 citations
- Multimodal Generative Learning Utilizing Jensen-Shannon-DivergenceThomas M. Sutter, Imant Daunhawer, Julia E. VogtNeurIPS 2020 · 105 citations
- Capturing Label Characteristics in VAEsTom Joy, Sebastian M. Schmon, Philip H. S. Torr, Siddharth Narayanaswamy et al.ICLR 2021 · 54 citations
- Relating by Contrasting: A Data-efficient Framework for Multimodal Generative ModelsYuge Shi, Brooks Paige, Philip H. S. Torr, N. SiddharthICLR 2021 · 42 citations
Related papers
- A Span-based Multimodal Variational Autoencoder for Semi-supervised Multimodal Named Entity RecognitionBaohang Zhou, Ying Zhang, Kehui Song, Wenya Guo et al.EMNLP 2022 · 15 citations
- Multimodal Gaussian Mixture Variational Autoencoder with Consistency RegularizationsYarui Chen, Lehan Hong, Jianlin Shao, Jianning Yang et al.AAAI 2026
- Associative Variational Auto-Encoder with Distributed Latent Spaces and AssociatorsDae Ung Jo, Byeongju Lee, Jongwon Choi, Haanju Yoo et al.AAAI 2020 · 8 citations
- Incomplete Multi-View Multi-label Learning via Disentangled Representation and Label Semantic EmbeddingXu Yan, Jun Yin, Jie WenCVPR 2025
- Disentanglement of Variations with Multimodal Generative ModelingYijie Zhang, Yiyang Shen, Weiran WangICLR 2026 · 6 citations
