Characterizing the Predictive Impact of Modalities with Supervised Latent-Variable Modeling
Divyam Madaan, Sumit Chopra, Kyunghyun Cho
摘要
Despite the recent success of Multimodal Large Language Models (MLLMs), existing approaches predominantly assume the availability of multiple modalities during training and inference. In practice, multimodal data is often incomplete because modalities may be missing, collected asynchronously, or available only for a subset of examples. In this work, we propose PRIMO, a supervised latent-variable imputation model that quantifies the predictive impact of any missing modality within the multimodal learning setting. PRIMO enables the use of all available training examples, whether modalities are complete or partial. Specifically, it models the missing modality through a latent variable that captures its relationship with the observed modality in the context of prediction. During inference, we draw many samples from the learned distribution over the missing modality to both obtain the marginal predictive distribution (for the purpose of prediction) and analyze the impact of the missing modalities on the prediction for each instance. We evaluate PRIMO on a synthetic XOR dataset, Audio-Vision MNIST, and MIMIC-III for mortality and ICD-9 prediction. Across all datasets, PRIMO obtains performance comparable to unimodal baselines when a modality is fully missing and to multimodal baselines when all modalities are available. PRIMO quantifies the predictive impact of a modality at the instance level using a variance-based metric computed from predictions across latent completions. We visually demonstrate how varying completions of the missing modality result in a set of plausible labels.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper21
- MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding BenchmarkXiang Yue, Tianyu Zheng, Yuansheng Ni, Yubo Wang 等ACL 2025 · 被引用 377 次
- MMMU: A Massive Multi-Discipline Multimodal Understanding and Reasoning Benchmark for Expert AGIXiang Yue, Yuansheng Ni, Tianyu Zheng, Kai Zhang 等CVPR 2024 · 被引用 213 次
- Generalized Multimodal ELBOThomas M. Sutter, Imant Daunhawer, Julia E. VogtICLR 2021 · 被引用 130 次
- Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMsShengbang Tong, Zhuang Liu, Yuexiang Zhai, Yi Ma 等CVPR 2024 · 被引用 111 次
- Multimodal Generative Learning Utilizing Jensen-Shannon-DivergenceThomas M. Sutter, Imant Daunhawer, Julia E. VogtNeurIPS 2020 · 被引用 105 次
相关 Paper
- Distilled Prompt Learning for Incomplete Multimodal Survival PredictionYingxue Xu, Fengtao Zhou, Chenyu Zhao, Yihui Wang 等CVPR 2025
- Deep Correlated Prompting for Visual Recognition with Missing ModalitiesLianyu Hu, Tongkai Shi, Wei Feng, Fanhua Shang 等NeurIPS 2024 · 被引用 37 次
- M3Care: Learning with Missing Modalities in Multimodal Healthcare DataChaohe Zhang, Xu Chu, Liantao Ma, Yinghao Zhu 等KDD 2022 · 被引用 78 次
- LIMSSR: LLM-Driven Sequence-to-Score Reasoning under Training-Time Incomplete Multimodal ObservationsHuangbiao Xu, huanqi wu, Xiao Ke, Yuxin PengICML 2026 · 被引用 1 次
- Causal Representation Learning from Multimodal Clinical Records under Non-Random Modality MissingnessZihan Liang, Ziwen Pan, Ruoxuan XiongEMNLP 2025
