MMoE: Enhancing Multimodal Models with Mixtures of Multimodal Interaction Experts
Haofei Yu, Zhengyang Qi, Lawrence Jang, Russ Salakhutdinov, Louis-Philippe Morency, Paul Pu Liang
Abstract
Advances in multimodal models have greatly improved how interactions relevant to various tasks are modeled. Today’s multimodal models mainly focus on the correspondence between images and text, using this for tasks like image-text matching. However, this covers only a subset of real-world interactions. Novel interactions, such as sarcasm expressed through opposing spoken words and gestures or humor expressed through utterances and tone of voice, remain challenging. In this paper, we introduce an approach to enhance multimodal models, which we call Multimodal Mixtures of Experts (MMoE). The key idea in MMoE is to train separate expert models for each type of multimodal interaction, such as redundancy present in both modalities, uniqueness in one modality, or synergy that emerges when both modalities are fused. On a sarcasm detection task (MUStARD) and a humor detection task (URFUNNY), we obtain new state-of-the-art results. MMoE is also able to be applied to various types of models to gain improvement.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers8
- MAESTRO : Adaptive Sparse Attention and Robust Learning for Multimodal Dynamic Time SeriesPayal Mohapatra, Yueyuan Sui, Akash Pandey, Stephen Xia et al.NeurIPS 2025 · 19 citations
- UniT: Unified Multimodal Chain-of-Thought Test-time ScalingLeon Liangyu Chen, Haoyu Ma, Zhipeng Fan, Ziqi Huang et al.CVPR 2026 · 7 citations
- REDEEMing Modality Information Loss: Retrieval-Guided Conditional Generation for Severely Modality Missing LearningJian Lang, Rongpei Hong, Zhangtao Cheng, Ting Zhong et al.KDD 2025 · 5 citations
- Compose and Fuse: Revisiting the Foundational Bottlenecks in Multimodal ReasoningYucheng Wang, Yifan Hou, Aydin Javadov, Mubashara Akhtar et al.ICLR 2026 · 3 citations
- SatireDecoder: Visual Cascaded Decoupling for Enhancing Satirical Image ComprehensionYue Jiang, Haiwei Xue, Minghao Han, Mingcheng Li et al.AAAI 2026 · 2 citations
Builds on25
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
Related papers
- Sentiment and Emotion help Sarcasm? A Multi-task Learning Framework for Multi-Modal Sarcasm, Sentiment and Emotion AnalysisDushyant Singh Chauhan, Dhanush S. R, Asif Ekbal, Pushpak BhattacharyyaACL 2020 · 131 citations
- Conditional Information Bottleneck for Multimodal Fusion: Overcoming Shortcut Learning in Sarcasm DetectionYihua Wang, Qi Jia, Cong Xu, Feiyu Chen et al.AAAI 2026
- Nice Perfume. How Long Did You Marinate in It? Multimodal Sarcasm ExplanationPoorav Desai, Tanmoy Chakraborty, Md. Shad AkhtarAAAI 2022 · 49 citations
- Predict and Use: Harnessing Predicted Gaze to Improve Multimodal Sarcasm DetectionDivyank Tiwari, Diptesh Kanojia, Anupama Ray, Apoorva Nunna et al.EMNLP 2023 · 12 citations
- Mutual-Enhanced Incongruity Learning Network for Multi-Modal Sarcasm DetectionYang Qiao, Liqiang Jing, Xuemeng Song, Xiaolin Chen et al.AAAI 2023 · 84 citations
