D2R: Dual-Branch Dynamic Routing Network for Multimodal Sentiment Detection
Yifan Chen, Kuntao Li, Weixing Mai, Qiaofeng Wu, Yun Xue, Fenghuan Li
Abstract
Multimodal sentiment detection aims to classify the sentiment polarity of a given imagetext pair. Existing approaches apply the same fixed framework to all input samples, lacking the flexibility to adapt to different image-text pairs. Furthermore, the interaction patterns of these methods are overly homogenized, limiting the model's capacity to extract multimodal sentiment information effectively. In this paper, we develop a Dual-Branch Dynamic Routing Network (D 2 R), which is the first multimodal dynamic interaction model towards multimodal sentiment detection. Specifically, we design six independent units to simulate inter-and intramodal information interactions without depending on any existing fixed frameworks. Additionally, we configure a soft router in each unit to guide path generation and introduce the path regularization term to optimize these inference paths. Comprehensive experiments on three publicly available datasets demonstrate the superiority of our proposed model over state-ofthe-art methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7e9db761-66dc-4c15-aebf-b1c2cd361e7dCited by top-tier papers4
- CLCR: Cross-Level Semantic Collaborative Representation for Multimodal LearningChunlei Meng, Guanhong Huang, Rong Fu, Runmin Jian et al.CVPR 2026 · 9 citations
- Tri-Subspaces Disentanglement for Multimodal Sentiment AnalysisChunlei Meng, Jiabin Luo, Zhenglin Yan, Zhenyu Yu et al.CVPR 2026 · 7 citations
- Group Cognition Learning: Making Everything Better Through Controlled Two-Stage Agents CollaborationChunlei Meng, Pengbin Feng, Rong Fu, Hoi Leong Lee et al.ICML 2026 · 2 citations
- Notes-guided MLLM Reasoning: Enhancing MLLM with Knowledge and Visual Notes for Visual Question AnsweringWenlong Fang, Qiaofeng Wu, Jing Chen, Yun XueCVPR 2025
Builds on12
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Similarity Reasoning and Filtration for Image-Text MatchingHaiwen Diao, Ying Zhang, Lin Ma, Huchuan LuAAAI 2021 · 413 citations
- Dynamic Modality Interaction Modeling for Image-Text RetrievalLeigang Qu, Meng Liu, Jianlong Wu, Zan Gao et al.SIGIR 2021 · 187 citations
- Reasoning with Multimodal Sarcastic Tweets via Modeling Cross-Modality Contrast and Semantic AssociationNan Xu, Zhixiong Zeng, Wenji MaoACL 2020 · 153 citations
- TRAR: Routing the Attention Spans in Transformer for Visual Question AnsweringYiyi Zhou, Tianhe Ren, Chaoyang Zhu, Xiaoshuai Sun et al.ICCV 2021 · 128 citations
Related papers
- Dynamic Routing Transformer Network for Multimodal Sarcasm DetectionYuan Tian, Nan Xu, Ruike Zhang, Wenji MaoACL 2023 · 40 citations
- DRDF: Determining the Importance of Different Multimodal Information with Dual-Router Dynamic FrameworkHaiwen Hong, Xuan Jin, Yin Zhang, Yunqing Hu et al.ACM MM 2021
- Seek Common Ground While Reserving Differences: Semi-Supervised Image-Text Sentiment RecognitionWuyou Xia, Guoli Jia, Sicheng Zhao, Jufeng YangCVPR 2025
- Robust Multimodal Sentiment Analysis of Image-Text Pairs by Distribution-Based Feature Recovery and FusionDaiqing Wu, Dongbao Yang, Yu Zhou, Can MaACM MM 2024 · 13 citations
- Multimodal Graph Representation Learning with Dynamic Information PathwaysXiaobin Hong, Mingkai Lin, Xiaoli Wang, Chaoqun Wang et al.AAAI 2026 · 1 citation
