MultiMoDN - Multimodal, Multi-Task, Interpretable Modular Networks
Vinitra Swamy, Malika Satayeva, Jibril Frej, Thierry Bossy, Thijs Vogels, Martin Jaggi, Tanja Käser, Mary-Anne Hartley
摘要
Predicting multiple real-world tasks in a single model often requires a particularly diverse feature space. Multimodal (MM) models aim to extract the synergistic predictive potential of multiple data types to create a shared feature space with aligned semantic meaning across inputs of drastically varying sizes (i.e. images, text, sound). Most current MM architectures fuse these representations in parallel, which not only limits their interpretability but also creates a dependency on modality availability. We present MultiModN, a multimodal, modular network that fuses latent representations in a sequence of any number, combination, or type of modality while providing granular real-time predictive feedback on any number or combination of predictive tasks. MultiModN's composable pipeline is interpretable-by-design, as well as innately multi-task and robust to the fundamental issue of biased missingness. We perform four experiments on several benchmark MM datasets across 10 real-world tasks (predicting medical diagnoses, academic performance, and weather), and show that MultiModN's sequential MM fusion does not compromise performance compared with a baseline of parallel fusion. By simulating the challenging bias of missing not-at-random (MNAR), this work shows that, contrary to MultiModN, parallel fusion baselines erroneously learn MNAR and suffer catastrophic failure when faced with different patterns of MNAR at inference. To the best of our knowledge, this is the first inherently MNAR-resistant approach to MM modeling. In conclusion, MultiModN provides granular insights, robustness, and flexibility without compromising performance. * denotes equal contribution 37th Conference on Neural Information Processing Systems (NeurIPS 2023).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- HEALNet: Multimodal Fusion for Heterogeneous Biomedical DataKonstantin Hemker, Nikola Simidjievski, Mateja JamnikNeurIPS 2024 · 被引用 80 次
- Mixture of Cognitive Reasoners: Modular Reasoning with Brain-Like SpecializationBadr AlKhamissi, C. Nicolò De Sabbata, Greta Tuckute, Zeming Chen 等ICLR 2026 · 被引用 12 次
- Disentangled Cross-Modal Representation Learning with Enhanced Mutual SupervisionLu Gao, Wenlan Chen, Daoyuan Wang, Fei Guo 等NeurIPS 2025 · 被引用 5 次
- FlexCare: Leveraging Cross-Task Synergy for Flexible Multimodal Healthcare PredictionMuhao Xu, Zhenfeng Zhu, Youru Li, Shuai Zheng 等KDD 2024 · 被引用 5 次
- Sparse CLIP: Co-Optimizing Interpretability and Performance in Contrastive LearningChuan Qin, Constantin Venhoff, Sonia Joseph, Fanyi Xiao 等ICLR 2026 · 被引用 4 次
它引用的顶会 Paper4
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech 等NeurIPS 2022 · 被引用 6,707 次
- UniT: Multimodal Multitask Learning with a Unified TransformerRonghang Hu, Amanpreet SinghICCV 2021 · 被引用 354 次
- Are Multimodal Transformers Robust to Missing Modality?Mengmeng Ma, Jian Ren, Long Zhao, Davide Testuggine 等CVPR 2022 · 被引用 153 次
- MultiViz: Towards Visualizing and Understanding Multimodal ModelsPaul Pu Liang, Yiwei Lyu, Gunjan Chhablani, Nihal Jain 等ICLR 2023 · 被引用 15 次
相关 Paper
- Missing Modality Imagination Network for Emotion Recognition with Uncertain Missing ModalitiesJinming Zhao, Ruichen Li, Qin JinACL 2021
- Scalable Medical Multimodal Fusion via Symmetric Consistency ModelingXiaowen Sun, Hui Liu, Gongguan Chen, Ning MaoICML 2026
- Towards Good Practices for Missing Modality Robust Action RecognitionSangmin Woo, Sumin Lee, Yeonju Park, Muhammad Adi Nugroho 等AAAI 2023 · 被引用 80 次
- Compose and Fuse: Revisiting the Foundational Bottlenecks in Multimodal ReasoningYucheng Wang, Yifan Hou, Aydin Javadov, Mubashara Akhtar 等ICLR 2026 · 被引用 3 次
- OmniField: Conditioned Neural Fields for Robust Multimodal Spatiotemporal LearningKevin Valencia, Thilina Balasooriya, Xihaier Luo, Shinjae Yoo 等ICLR 2026 · 被引用 1 次
