Building Massively Multimodal Foundation Models with Interaction-aware Mixture-of-Experts
Xing Han, Hsing-Huan Chung, Joydeep Ghosh, Paul Pu Liang, Suchi Saria
Abstract
Modern applications increasingly involve many heterogeneous input streams, such as clinical sensors, wearable device data, imaging, and text, each with distinct measurement models, sampling rates, and noise characteristics. We define this as massively multimodal setting, where each sensor constitutes a separate modality. As modality counts grow, capturing their complex, time-varying interactions such as delayed physiological cascades between sensors, has becomes essential yet challenging. Mixture-of-Experts (MoE) architectures are naturally suited for this setting since their sparse routing mechanism enables efficient scaling across many modalities. However, existing MoE architectures route tokens based on similarity alone, overlooking the rich temporal dependencies across modalities: this prevents the model from capturing delayed cross-modal effects, leading to suboptimal expert specialization and reduced accuracy. We propose a framework that explicitly quantifies temporal dependencies between modality pairs across multiple discrete time intervals, defined as delays between an event in one input stream and its manifested effect in another, and uses these to guide MoE routing. A interaction-aware router dispatches tokens to specialized experts based on interaction type. This principled routing enables experts to learn generalizable interaction-processing skills. Experiments across healthcare, activity recognition, and affective computing benchmarks demonstrate substantial performance gains and interpretable routing patterns aligned with domain knowledge.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8d5c9ef2-58c3-408c-9914-eac046b0c561Cited by top-tier papers1
Ask how each one uses itBuilds on19
- Scaling Vision with Sparse Mixture of ExpertsCarlos Riquelme, Joan Puigcerver, Basil Mustafa, Maxim Neumann et al.NeurIPS 2021 · 1,213 citations
- Integrating Multimodal Information in Large Pretrained TransformersWasifur Rahman, Md. Kamrul Hasan, Sangwu Lee, AmirAli Bagher Zadeh et al.ACL 2020 · 584 citations
- Multimodal Contrastive Learning with LIMoE: the Language-Image Mixture of ExpertsBasil Mustafa, Carlos Riquelme, Joan Puigcerver, Rodolphe Jenatton et al.NeurIPS 2022 · 359 citations
- Multi-Time Attention Networks for Irregularly Sampled Time SeriesSatya Narayan Shukla, Benjamin M. MarlinICLR 2021 · 301 citations
- FuseMoE: Mixture-of-Experts Transformers for Fleximodal FusionXing Han, Huy Nguyen, Carl Harris, Nhat Ho et al.NeurIPS 2024 · 129 citations
Related papers
- I2MoE: Interpretable Multimodal Interaction-aware Mixture-of-ExpertsJiayi Xin, Sukwon Yun, Jie Peng, Inyoung Choi et al.ICML 2025
- Flex-MoE: Modeling Arbitrary Modality Combination via the Flexible Mixture-of-ExpertsSukwon Yun, Inyoung Choi, Jie Peng, Yangfan Wu et al.NeurIPS 2024 · 98 citations
- Soft Modality-Guided Expert Specialization in MoE-VLMsZi-Hao Bo, Yaqian Li, Anzhou Hou, Rinyoichi Takezoe et al.CVPR 2026
- How Many Experts Are Enough? Towards Optimal Semantic Specialization for Mixture-of-ExpertsSumin Park, Noseong ParkAAAI 2026
- MAESTRO : Adaptive Sparse Attention and Robust Learning for Multimodal Dynamic Time SeriesPayal Mohapatra, Yueyuan Sui, Akash Pandey, Stephen Xia et al.NeurIPS 2025 · 19 citations
