Multimodal Routing: Improving Local and Global Interpretability of Multimodal Language Analysis
Yao-Hung Hubert Tsai, Martin Q. Ma, Muqiao Yang, Ruslan Salakhutdinov, Louis-Philippe Morency
Abstract
The human language can be expressed through multiple sources of information known as modalities, including tones of voice, facial gestures, and spoken language. Recent multimodal learning with strong performances on human-centric tasks such as sentiment analysis and emotion recognition are often blackbox, with very limited interpretability. In this paper we propose Multimodal Routing, which dynamically adjusts weights between input modalities and output representations differently for each input sample. Multimodal routing can identify relative importance of both individual modalities and cross-modality features. Moreover, the weight assignment by routing allows us to interpret modalityprediction relationships not only globally (i.e. general trends over the whole dataset), but also locally for each single input sample, meanwhile keeping competitive performance compared to state-of-the-art methods. * indicates equal contribution. Code is available at https://github.com/martinmamql/ multimodal_routing .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5cdf5245-4ece-44dc-a8d4-b00bfb74c785Cited by top-tier papers9
- Learning Modality-Specific and -Agnostic Representations for Asynchronous Multimodal Language SequencesDingkang Yang, Haopeng Kuang, Shuai Huang, Lihua ZhangACM MM 2022 · 64 citations
- Joyful: Joint Modality Fusion and Graph Contrastive Learning for Multimoda Emotion RecognitionDongyuan Li, Yusong Wang, Kotaro Funakoshi, Manabu OkumuraEMNLP 2023 · 46 citations
- Curriculum Learning Meets Weakly Supervised Multimodal Correlation LearningSijie Mai, Ya Sun, Haifeng HuEMNLP 2022 · 9 citations
- Improving Multimodal Fusion with Hierarchical Mutual Information Maximization for Multimodal Sentiment AnalysisWei Han, Hui Chen, Soujanya PoriaEMNLP 2021 · 9 citations
- Towards Explainable Fusion and Balanced Learning in Multimodal Sentiment AnalysisMiaosen Luo, Yuncheng Jiang, Sijie MaiACM MM 2025 · 7 citations
Related papers
- D2R: Dual-Branch Dynamic Routing Network for Multimodal Sentiment DetectionYifan Chen, Kuntao Li, Weixing Mai, Qiaofeng Wu et al.EMNLP 2024 · 10 citations
- Enhanced Experts with Uncertainty-Aware Routing for Multimodal Sentiment AnalysisZixian Gao, Disen Hu, Xun Jiang, Huimin Lu et al.ACM MM 2024 · 17 citations
- M2Lens: Visualizing and Explaining Multimodal Models for Sentiment AnalysisXingbo Wang, Jianben He, Zhihua Jin, Muqiao Yang et al.IEEE VIS 2021 · 5 citations
- Conversation Understanding using Relational Temporal Graph Neural Networks with Auxiliary Cross-Modality InteractionCam-Van Thi Nguyen, Anh-Tuan Mai, The-Son Le, Hai-Dang Kieu et al.EMNLP 2023 · 34 citations
- Learning to Route: Per-Sample Adaptive Routing for Multimodal Multitask PredictionMarzieh Ajirak, Oded Bein, Ellen Rose Bowen, Dora Kanellopoulos et al.NeurIPS 2025 · 1 citation
