SimMLM: A Simple Framework for Multi-Modal Learning with Missing Modality
Sijie Li, Chen Chen, Jungong Han
Abstract
In this paper, we propose SimMLM, a simple yet powerful framework for multimodal learning with missing modalities. Unlike existing approaches that rely on sophisticated network architectures or complex data imputation techniques, SimMLM provides a generic and effective solution that can adapt to various missing modality scenarios with improved accuracy and robustness. Specifically, SimMLM consists of a generic Dynamic Mixture of Modality Experts (DMoME) architecture, featuring a dynamic, learnable gating mechanism that automatically adjusts each modality's contribution in both full and partial modality settings. A key innovation of SimMLM is the proposed More vs. Fewer (MoFe) ranking loss, which ensures that task accuracy improves or remains stable as more modalities are made available. This aligns the model with an intuitive principle: removing one or more modalities should not increase accuracy. We validate SimMLM on multimodal medical image segmentation (BraTS 2018) and multimodal classification (UPMC Food-101, avMNIST) tasks, where it consistently surpasses competitive methods, demonstrating superior accuracy, interpretability, robustness, and reliability across both complete and missing modality scenarios at test time.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ea5e750f-2714-449a-822b-a553db086405Cited by top-tier papers3
- Inference-Time Dynamic Modality Selection for Incomplete Multimodal ClassificationSiyi Du, Xinzhe Luo, Declan O'regan, Chen QinICLR 2026 · 4 citations
- Retrieving to Recover: Towards Incomplete Audio-Visual Question Answering via Semantic-consistent PurificationJiayu Zhang, Shuo Ye, Qilang Ye, Zihan Song et al.ACL 2026 · 2 citations
- AOEPT: Breaking the Implicit Modality-Reduction Bottleneck in Modality-Missing Prompt TuningJian Lang, Hong, Ting Zhong, Fan ZhouICML 2026 · 1 citation
Builds on19
- GShard: Scaling Giant Models with Conditional Computation and Automatic ShardingDmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen et al.ICLR 2021 · 1,954 citations
- Mind the Gap: Understanding the Modality Gap in Multi-modal Contrastive Representation LearningWeixin Liang, Yuhui Zhang, Yongchan Kwon, Serena Yeung et al.NeurIPS 2022 · 834 citations
- SMIL: Multimodal Learning with Severely Missing ModalityMengmeng Ma, Jian Ren, Long Zhao, Sergey Tulyakov et al.AAAI 2021 · 393 citations
- Incomplete Multimodality-Diffused Emotion RecognitionYuanzhi Wang, Yong Li, Zhen CuiNeurIPS 2023 · 155 citations
- Are Multimodal Transformers Robust to Missing Modality?Mengmeng Ma, Jian Ren, Long Zhao, Davide Testuggine et al.CVPR 2022 · 153 citations
Related papers
- Taming Cascaded Mixture-of-Experts for Modality-missing Multi-modal Salient Object DetectionKunpeng Wang, Feifan Sun, Keke ChenAAAI 2026
- Flex-MoE: Modeling Arbitrary Modality Combination via the Flexible Mixture-of-ExpertsSukwon Yun, Inyoung Choi, Jie Peng, Yangfan Wu et al.NeurIPS 2024 · 98 citations
- Gradient-Guided Modality Decoupling for Missing-Modality RobustnessHao Wang, Shengda Luo, Guosheng Hu, Jianguo ZhangAAAI 2024 · 20 citations
- Rethinking Gating Mechanism in Sparse MoE: Handling Arbitrary Modality Inputs with Confidence-Guided GateLiangwei Zheng, Wei Emma Zhang, Mingyu Guo, Olaf Maennel et al.ICML 2026 · 7 citations
- Auto-GAN: Self-Supervised Collaborative Learning for Medical Image SynthesisBing Cao, Han Zhang, Nannan Wang, Xinbo Gao et al.AAAI 2020 · 94 citations
