DRDF: Determining the Importance of Different Multimodal Information with Dual-Router Dynamic Framework
Haiwen Hong, Xuan Jin, Yin Zhang, Yunqing Hu, Jingfeng Zhang, Yuan He, Hui Xue
Abstract
In multimodal tasks, the importance of text and image modal information often varies for different input cases. To model the difference of importance of different modal information, we propose a high-performance and highly general Dual-Router Dynamic Framework (DRDF), consisting of Dual-Router, MWF-Layer, experts and expert fusion unit. The text router and image router in Dual-Router take text modal information and image modal information respectively, and MWF-Layer is responsible to determine the importance of modal information. Based on the result of the determination, MWF-Layer generates fused weights for the subsequent experts fusion. Experts can adopt a variety of backbones that match the current multimodal or unimodal task. DRDF features high generality and modularity, and we test 12 backbones such as Visual BERT and their corresponding DRDF instances on the multimodal dataset Hateful memes, and unimodal datasets CIFAR10, CIFAR100, and TinyImagenet. Our DRDF instance outperforms those backbones. We also validate the effectiveness of components of DRDF by ablation studies, and discuss the reasons and ideas of DRDF design.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0a9b7fc5-7915-4faf-9874-d86e52d3562bBuilds on7
- VL-BERT: Pre-training of Generic Visual-Linguistic RepresentationsWeijie Su, Xizhou Zhu, Yue Cao, Bin Li et al.ICLR 2020 · 1,825 citations
- The Hateful Memes Challenge: Detecting Hate Speech in Multimodal MemesDouwe Kiela, Hamed Firooz, Aravind Mohan, Vedanuj Goswami et al.NeurIPS 2020 · 1,022 citations
- Unicoder-VL: A Universal Encoder for Vision and Language by Cross-Modal Pre-TrainingGen Li, Nan Duan, Yuejian Fang, Ming Gong et al.AAAI 2020 · 966 citations
- Deformable Kernels: Adapting Effective Receptive Fields for Object DeformationHang Gao, Xizhou Zhu, Stephen Lin, Jifeng DaiICLR 2020 · 72 citations
- NBDT: Neural-Backed Decision TreeAlvin Wan, Lisa Dunlap, Daniel Ho, Jihan Yin et al.ICLR 2021 · 56 citations
Related papers
- D2R: Dual-Branch Dynamic Routing Network for Multimodal Sentiment DetectionYifan Chen, Kuntao Li, Weixing Mai, Qiaofeng Wu et al.EMNLP 2024 · 10 citations
- Dynamic Modality Interaction Modeling for Image-Text RetrievalLeigang Qu, Meng Liu, Jianlong Wu, Zan Gao et al.SIGIR 2021 · 187 citations
- Dynamic Routing Transformer Network for Multimodal Sarcasm DetectionYuan Tian, Nan Xu, Ruike Zhang, Wenji MaoACL 2023 · 40 citations
- Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical FindingsQiong Wu, Wenhao Lin, Yiyi Zhou, Weihao Ye et al.NeurIPS 2025 · 16 citations
- Expanding Large Pre-trained Unimodal Models with Multimodal Information Injection for Image-Text Multimodal ClassificationTao Liang, Guosheng Lin, Mingyang Wan, Tianrui Li et al.CVPR 2022 · 39 citations
