Mitigating Modality Collapse in Multimodal VAEs via Impartial Optimization
Adrián Javaloy, Maryam Meghdadi, Isabel Valera
Abstract
A number of variational autoencoders (VAEs) have recently emerged with the aim of modeling multimodal data, e.g., to jointly model images and their corresponding captions. Still, multimodal VAEs tend to focus solely on a subset of the modalities, e.g., by fitting the image while neglecting the caption. We refer to this limitation as modality collapse. In this work, we argue that this effect is a consequence of conflicting gradients during multimodal VAE training. We show how to detect the sub-graphs in the computational graphs where gradients conflict (impartiality blocks), as well as how to leverage existing gradient-conflict solutions from multitask learning to mitigate modality collapse. That is, to ensure impartial optimization across modalities. We apply our training framework to several multimodal VAE models, losses and datasets from the literature, and empirically show that our framework significantly improves the reconstruction performance, conditional generation, and coherence of the latent space across modalities.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a7f9377b-5e1f-4e5a-b224-7920ea5e369eCited by top-tier papers14
- Multimodal Patient Representation Learning with Missing Modalities and LabelsZhenbang Wu, Anant Dadu, Nicholas J. Tustison, Brian B. Avants et al.ICLR 2024 · 38 citations
- Intra- and Inter-Modal Curriculum for Multimodal LearningYuwei Zhou, Xin Wang, Hong Chen, Xuguang Duan et al.ACM MM 2023 · 28 citations
- Cooperation in the Latent Space: The Benefits of Adding Mixture Components in Variational AutoencodersOskar Kviman, Ricky Molén, Alexandra Hotti, Semih Kurt et al.ICML 2023 · 16 citations
- Curriculum-Listener: Consistency- and Complementarity-Aware Audio-Enhanced Temporal Sentence GroundingHoulun Chen, Xin Wang, Xiaohan Lan, Hong Chen et al.ACM MM 2023 · 13 citations
- Ada2I: Enhancing Modality Balance for Multimodal Conversational Emotion RecognitionCam-Van Thi Nguyen, The-Son Le, Anh-Tuan Mai, Duc-Trong LeACM MM 2024 · 13 citations
Builds on9
- Gradient Surgery for Multi-Task LearningTianhe Yu, Saurabh Kumar, Abhishek Gupta, Sergey Levine et al.NeurIPS 2020 · 2,261 citations
- Conflict-Averse Gradient Descent for Multi-task learningBo Liu, Xingchao Liu, Xiaojie Jin, Peter Stone et al.NeurIPS 2021 · 686 citations
- Just Pick a Sign: Optimizing Deep Multitask Models with Gradient Sign DropoutZhao Chen, Jiquan Ngiam, Yanping Huang, Thang Luong et al.NeurIPS 2020 · 313 citations
- From Variational to Deterministic AutoencodersPartha Ghosh, Mehdi S. M. Sajjadi, Antonio Vergari, Michael J. Black et al.ICLR 2020 · 298 citations
- Towards Impartial Multi-task LearningLiyang Liu, Yi Li, Zhanghui Kuang, Jing-Hao Xue et al.ICLR 2021 · 228 citations
Related papers
- Multimodal Fusion via Self-Consistent Task-Gradient FieldsJiayu Xiong, Jing Wang, Jun Xue, Wanlong Wang et al.ICML 2026
- Relating by Contrasting: A Data-efficient Framework for Multimodal Generative ModelsYuge Shi, Brooks Paige, Philip H. S. Torr, N. SiddharthICLR 2021 · 42 citations
- Unity by Diversity: Improved Representation Learning for Multimodal VAEsThomas M. Sutter, Yang Meng, Andrea Agostini, Daphné Chopard et al.NeurIPS 2024 · 21 citations
- Incomplete Cross-modal Retrieval with Dual-Aligned Variational AutoencodersMengmeng Jing, Jingjing Li, Lei Zhu, Ke Lu et al.ACM MM 2020 · 63 citations
- On the Limitations of Multimodal VAEsImant Daunhawer, Thomas M. Sutter, Kieran Chin-Cheong, Emanuele Palumbo et al.ICLR 2022 · 50 citations
