An Optimal Transport-based Latent Mixer for Robust Multi-modal Learning
Fengjiao Gong, Angxiao Yue, Hongteng Xu
摘要
Multi-modal learning aims to learn predictive models based on the data from different modalities. However, due to the requirement of data security and privacy protection, real-world multi-modal data are often scattered to different agents and cannot be shared across the agents, which limits the application of existing multi-modal learning methods. To achieve robust multi-modal learning in such a challenging scenario, we propose a novel optimal transport-based mixer (OTM), which works as an effective latent code alignment and augmentation method for unaligned and distributed multi-modal data. In particular, we train a Wasserstein autoencoder (WAE) for each agent, which encodes its single modal samples in a latent space. Through a central server, the proposed OTM computes a stochastic fused Gromov-Wasserstein barycenter (FGWB) to mix different modalities' latent codes, so that each agent applies the barycenter to reconstruct its samples. This method neither requires well-aligned multi-modal data nor assumes the data to share the same latent distribution, and each agent can learn a specific model based on multi-modal data while achieving inference based on its local modality. Experiments on multi-modal clustering and classification demonstrate that the models learned with the OTM method outperform the corresponding baselines.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper10
- CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesSangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh 等ICCV 2019 · 被引用 5,843 次
- SMIL: Multimodal Learning with Severely Missing ModalityMengmeng Ma, Jian Ren, Long Zhao, Sergey Tulyakov 等AAAI 2021 · 被引用 393 次
- Multiplicative Interactions and Where to Find ThemSiddhant M. Jayakumar, Wojciech M. Czarnecki, Jacob Menick, Jonathan Schwarz 等ICLR 2020 · 被引用 152 次
- i-Mix: A Domain-Agnostic Strategy for Contrastive Representation LearningKibok Lee, Yian Zhu, Kihyuk Sohn, Chun-Liang Li 等ICLR 2021 · 被引用 133 次
- Dynamically Adjust Word Representations Using Unaligned Multimodal InformationJiwei Guo, Jiajia Tang, Weichen Dai, Yu Ding 等ACM MM 2022 · 被引用 59 次
相关 Paper
- A Wasserstein Minimax Framework for Mixed Linear RegressionTheo Diamandis, Yonina C. Eldar, Alireza Fallah, Farzan Farnia 等ICML 2021 · 被引用 7 次
- Multimodal Variational Autoencoder: A Barycentric ViewPeijie Qiu, Wenhui Zhu, Sayantan Kumar, Xiwen Chen 等AAAI 2025 · 被引用 2 次
- MFC: Mixed Federated Clustering based on Cross-modal Feature DecouplingXiaxia He, Boyue Wang, Junbin Gao, Yongli Hu 等KDD 2026
- MoReL: Multi-omics Relational LearningArman Hasanzadeh, Ehsan Hajiramezanali, Nick Duffield, Xiaoning QianICLR 2022 · 被引用 7 次
- Prototype-guided Bilateral Alignment Multimodal Federated LearningTianchi Liao, Lele Fu, Sheng Huang, Qing Hu 等ICML 2026
