An Optimal Transport-based Latent Mixer for Robust Multi-modal Learning
Fengjiao Gong, Angxiao Yue, Hongteng Xu
Abstract
Multi-modal learning aims to learn predictive models based on the data from different modalities. However, due to the requirement of data security and privacy protection, real-world multi-modal data are often scattered to different agents and cannot be shared across the agents, which limits the application of existing multi-modal learning methods. To achieve robust multi-modal learning in such a challenging scenario, we propose a novel optimal transport-based mixer (OTM), which works as an effective latent code alignment and augmentation method for unaligned and distributed multi-modal data. In particular, we train a Wasserstein autoencoder (WAE) for each agent, which encodes its single modal samples in a latent space. Through a central server, the proposed OTM computes a stochastic fused Gromov-Wasserstein barycenter (FGWB) to mix different modalities' latent codes, so that each agent applies the barycenter to reconstruct its samples. This method neither requires well-aligned multi-modal data nor assumes the data to share the same latent distribution, and each agent can learn a specific model based on multi-modal data while achieving inference based on its local modality. Experiments on multi-modal clustering and classification demonstrate that the models learned with the OTM method outperform the corresponding baselines.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 10f18254-a832-4534-ab9f-9499cf5bf36aBuilds on10
- CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesSangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh et al.ICCV 2019 · 5,843 citations
- SMIL: Multimodal Learning with Severely Missing ModalityMengmeng Ma, Jian Ren, Long Zhao, Sergey Tulyakov et al.AAAI 2021 · 393 citations
- Multiplicative Interactions and Where to Find ThemSiddhant M. Jayakumar, Wojciech M. Czarnecki, Jacob Menick, Jonathan Schwarz et al.ICLR 2020 · 152 citations
- i-Mix: A Domain-Agnostic Strategy for Contrastive Representation LearningKibok Lee, Yian Zhu, Kihyuk Sohn, Chun-Liang Li et al.ICLR 2021 · 133 citations
- Dynamically Adjust Word Representations Using Unaligned Multimodal InformationJiwei Guo, Jiajia Tang, Weichen Dai, Yu Ding et al.ACM MM 2022 · 59 citations
Related papers
- A Wasserstein Minimax Framework for Mixed Linear RegressionTheo Diamandis, Yonina C. Eldar, Alireza Fallah, Farzan Farnia et al.ICML 2021 · 7 citations
- Multimodal Variational Autoencoder: A Barycentric ViewPeijie Qiu, Wenhui Zhu, Sayantan Kumar, Xiwen Chen et al.AAAI 2025 · 2 citations
- MFC: Mixed Federated Clustering based on Cross-modal Feature DecouplingXiaxia He, Boyue Wang, Junbin Gao, Yongli Hu et al.KDD 2026
- MoReL: Multi-omics Relational LearningArman Hasanzadeh, Ehsan Hajiramezanali, Nick Duffield, Xiaoning QianICLR 2022 · 7 citations
- Prototype-guided Bilateral Alignment Multimodal Federated LearningTianchi Liao, Lele Fu, Sheng Huang, Qing Hu et al.ICML 2026
