Learning Unseen Modality Interaction
Yunhua Zhang, Hazel Doughty, Cees Snoek
Abstract
Multimodal learning assumes all modality combinations of interest are available during training to learn cross-modal correspondences. In this paper, we challenge this modality-complete assumption for multimodal learning and instead strive for generalization to unseen modality combinations during inference. We pose the problem of unseen modality interaction and introduce a first solution. It exploits a module that projects the multidimensional features of different modalities into a common space with rich information preserved. This allows the information to be accumulated with a simple summation operation across available modalities. To reduce overfitting to less discriminative modality combinations during training, we further improve the model learning with pseudo-supervision indicating the reliability of a modality's prediction. We demonstrate that our approach is effective for diverse tasks and modalities by evaluating it for multimodal video classification, robot state regression, and multimedia retrieval. Project website: https://xiaobai1217.github.io/Unseen-Modality-Interaction/ .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- Tackling Uncertain Correspondences for Multi-Modal Entity AlignmentLiyi Chen, Ying Sun, Shengzhe Zhang, Yuyang Ye et al.NeurIPS 2024 · 20 citations
- ReID5o: Achieving Omni Multi-modal Person Re-identification in a Single ModelJialong Zuo, Yongtai Deng, Mengdan Tan, Rui Jin et al.NeurIPS 2025 · 11 citations
- MoRA: Missing Modality Low-Rank Adaptation for Visual RecognitionShu Zhao, Nilesh A. Ahuja, Tan Yu, Tianyi Shen et al.ICLR 2026 · 5 citations
- Robust Modality-Incomplete Anomaly Detection: A Modality-Instructive Framework with BenchmarkBingchen Miao, Wenqiao Zhang, Juncheng Li, Wangyu Wu et al.ACM MM 2025 · 1 citation
- I2MoE: Interpretable Multimodal Interaction-aware Mixture-of-ExpertsJiayi Xin, Sukwon Yun, Jie Peng, Inyoung Choi et al.ICML 2025
Builds on25
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- Practical Secure Aggregation for Privacy-Preserving Machine LearningKallista A. Bonawitz, Vladimir Ivanov, Ben Kreuter, Antonio Marcedone et al.CCS 2017 · 3,936 citations
- Video Swin TransformerZe Liu, Jia Ning, Yue Cao, Yixuan Wei et al.CVPR 2022 · 1,847 citations
- Perceiver: General Perception with Iterative AttentionAndrew Jaegle, Felix Gimeno, Andy Brock, Oriol Vinyals et al.ICML 2021 · 1,399 citations
Related papers
- Achieving Cross Modal Generalization with Multimodal Unified RepresentationYan Xia, Hai Huang, Jieming Zhu, Zhou ZhaoNeurIPS 2023 · 84 citations
- On Uni-Modal Feature Learning in Supervised Multi-Modal LearningChenzhuang Du, Jiaye Teng, Tingle Li, Yichen Liu et al.ICML 2023 · 79 citations
- Towards Unified Vision-Language Models with Incomplete Multi-Modal InputsXiang Fang, Wanlong Fang, Changshuo Wang, Keke Tang et al.AAAI 2026 · 1 citation
- mDALU: Multi-Source Domain Adaptation and Label Unification with Partial DatasetsRui Gong, Dengxin Dai, Yuhua Chen, Wen Li et al.ICCV 2021 · 27 citations
- Towards Out-of-Modal Generalization without Instance-level Modal CorrespondenceZhuo Huang, Gang Niu, Bo Han, Masashi Sugiyama et al.ICLR 2025
