Relating by Contrasting: A Data-efficient Framework for Multimodal Generative Models
Yuge Shi, Brooks Paige, Philip H. S. Torr, N. Siddharth
Abstract
Multimodal learning for generative models often refers to the learning of abstract concepts from the commonality of information in multiple modalities, such as vision and language. While it has proven effective for learning generalisable representations, the training of such models often requires a large amount of "related" multimodal data that shares commonality, which can be expensive to come by. To mitigate this, we develop a novel contrastive framework for generative model learning, allowing us to train the model not just by the commonality between modalities, but by the distinction between "related" and "unrelated" multimodal data. We show in experiments that our method enables data-efficient multimodal learning on challenging datasets for various multimodal variational autoencoder (VAE) models. We also show that under our proposed framework, the generative model can accurately identify related samples from unrelated ones, making it possible to make use of the plentiful unlabeled, unpaired multimodal data.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 61b48276-9f33-470a-8dcd-e152fc494bb5Cited by top-tier papers15
- SMIL: Multimodal Learning with Severely Missing ModalityMengmeng Ma, Jian Ren, Long Zhao, Sergey Tulyakov et al.AAAI 2021 · 393 citations
- Generalized Multimodal ELBOThomas M. Sutter, Imant Daunhawer, Julia E. VogtICLR 2021 · 130 citations
- On the Limitations of Multimodal VAEsImant Daunhawer, Thomas M. Sutter, Kieran Chin-Cheong, Emanuele Palumbo et al.ICLR 2022 · 50 citations
- Mitigating Modality Collapse in Multimodal VAEs via Impartial OptimizationAdrián Javaloy, Maryam Meghdadi, Isabel ValeraICML 2022 · 49 citations
- Gaussian Mixture Variational Autoencoder with Contrastive Learning for Multi-Label ClassificationJunwen Bai, Shufeng Kong, Carla P. GomesICML 2022 · 48 citations
Builds on4
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Understanding Contrastive Representation Learning through Alignment and Uniformity on the HypersphereTongzhou Wang, Phillip IsolaICML 2020 · 2,360 citations
- Data-Efficient Image Recognition with Contrastive Predictive CodingOlivier J. HénaffICML 2020 · 1,553 citations
- Momentum Contrast for Unsupervised Visual Representation LearningKaiming He, Haoqi Fan, Yuxin Wu, Saining Xie et al.CVPR 2020
Related papers
- Multimodal Adversarially Learned Inference with Factorized DiscriminatorsWenxue Chen, Jianke ZhuAAAI 2022 · 3 citations
- Multimodal Gaussian Mixture Variational Autoencoder with Consistency RegularizationsYarui Chen, Lehan Hong, Jianlin Shao, Jianning Yang et al.AAAI 2026
- Associative Variational Auto-Encoder with Distributed Latent Spaces and AssociatorsDae Ung Jo, Byeongju Lee, Jongwon Choi, Haanju Yoo et al.AAAI 2020 · 8 citations
- Synthetic Data Can Also Teach: Synthesizing Effective Data for Unsupervised Visual Representation LearningYawen Wu, Zhepeng Wang, Dewen Zeng, Yiyu Shi et al.AAAI 2023 · 20 citations
- Connecting Multi-modal Contrastive RepresentationsZehan Wang, Yang Zhao, Xize Cheng, Haifeng Huang et al.NeurIPS 2023 · 60 citations
