Inference-Time Dynamic Modality Selection for Incomplete Multimodal Classification
Siyi Du, Xinzhe Luo, Declan O'regan, Chen Qin
Abstract
Multimodal deep learning (MDL) has achieved remarkable success across various domains, yet its practical deployment is often hindered by incomplete multimodal data. Existing incomplete MDL methods either discard missing modalities, risking the loss of valuable task-relevant information, or recover them, potentially introducing irrelevant noise, leading to the discarding-imputation dilemma. To address this dilemma, in this paper, we propose DyMo, a new inference-time dynamic modality selection framework that adaptively identifies and fuses reliable recovered modalities, fully exploring task-relevant information beyond the conventional discard-or-impute paradigm. Central to DyMo is a novel selection algorithm that maximizes multimodal task-relevant information for each test sample. Since direct estimation of such information at test time is intractable due to the unknown data distribution, we theoretically establish a connection between information and the task loss, which we compute at inference time as a tractable proxy. Building on this, a novel principled reward function is proposed to guide modality selection. In addition, we design a flexible multimodal network architecture compatible with arbitrary modality combinations, alongside a tailored training strategy for robust representation learning. Extensive experiments on diverse natural and medical image datasets show that DyMo significantly outperforms state-of-the-art incomplete/dynamic MDL methods across various missing-data scenarios. Our code is available at https://github.com//siyi-wind/DyMo .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext aa224739-34e0-4f2b-9c1b-7a85f1c8b652Builds on27
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- SMIL: Multimodal Learning with Severely Missing ModalityMengmeng Ma, Jian Ren, Long Zhao, Sergey Tulyakov et al.AAAI 2021 · 393 citations
- CoMatch: Semi-supervised Learning with Contrastive Graph RegularizationJunnan Li, Caiming Xiong, Steven C. H. HoiICCV 2021 · 333 citations
- Modality Competition: What Makes Joint Training of Multi-modal Network Fail in Deep Learning? (Provably)Yu Huang, Junyang Lin, Chang Zhou, Hongxia Yang et al.ICML 2022 · 168 citations
- Incomplete Multimodality-Diffused Emotion RecognitionYuanzhi Wang, Yong Li, Zhen CuiNeurIPS 2023 · 155 citations
Related papers
- MM-Align: Learning Optimal Transport-based Alignment Dynamics for Fast and Accurate Inference on Missing Modality SequencesWei Han, Hui Chen, Min-Yen Kan, Soujanya PoriaEMNLP 2022 · 13 citations
- SimMLM: A Simple Framework for Multi-Modal Learning with Missing ModalitySijie Li, Chen Chen, Jungong HanICCV 2025 · 14 citations
- CyIN: Cyclic Informative Latent Space for Bridging Complete and Incomplete Multimodal LearningRonghao Lin, Qiaolin He, Sijie Mai, Ying Zeng et al.NeurIPS 2025 · 7 citations
- Incomplete Modality Disentangled Representation for Ophthalmic Disease Grading and DiagnosisChengzhi Liu, Zile Huang, Zhe Chen, Feilong Tang et al.AAAI 2025 · 10 citations
- Missing Modality Imagination Network for Emotion Recognition with Uncertain Missing ModalitiesJinming Zhao, Ruichen Li, Qin JinACL 2021
