Adaptive Multimodal Fusion: Dynamic Attention Allocation for Intent Recognition
Bo Hu, Kai Zhang, Yanghai Zhang, Yuyang Ye
Abstract
In recent years, deep multimodal learning has seen significant advancements. However, there remains a lack of multimodal fusion methods capable of dynamically adjusting the weighting of information both within and across modalities based on input samples. In the domain of multimodal intent recognition, the text modality often contains the most relevant information for intent detection, while the audio and visual modalities provide comparatively less critical information. There is a significant variation in the density of important information across different modalities and samples. To address this challenge, we propose a Dynamic Attention Allocation Fusion (DAF) method with an adaptive network structure that dynamically allocates attention both within individual modalities and across multiple modalities. This approach enables the model to focus more effectively on the most informative modalities and their respective internal features. Furthermore, we introduce a multi-view contrastive learning framework based on DAF (MVCL-DAF). This framework uses distinct and isolated modules to process information from various modalities, taking inspiration from the way the human brain processes multimodal information. Each modality independently infers intent using its respective module, while DAF integrates the multimodal information to produce a comprehensive global intent prediction. The text modality, functioning as the primary modality due to its rich semantic content, guides the other modules in the multi-view contrastive learning process. Extensive experiments demonstrate that our approach significantly outperforms existing state-of-the-art methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3d31028c-4281-423d-94da-a494a0ee97f8Cited by top-tier papers4
- Evolutionary Multimodal Reasoning via Hierarchical Semantic Representation for Intent RecognitionQianrui Zhou, Hua Xu, Yunjin Gu, Yifan Wang et al.CVPR 2026 · 3 citations
- Who Should I Trust? Explicit Confidence-Focused Multimodal Intent RecognitionYi Liu, Qimeng Yang, Lanlan LuAAAI 2026
- Inconsistency-aware Multimodal Schrödinger Bridge for Deepfake LocalizationJiayu Xiong, Jing Wang, Qi Zhang, Wanlong Wang et al.CVPR 2026
- SeD-UD: An Influence-Driven and Hierarchically-Decoupled Information Bottleneck for Multimodal Intent RecognitionQin Li, Wenbo Zhang, Limei Liu, Han Peng et al.CVPR 2026
Builds on9
- Integrating Multimodal Information in Large Pretrained TransformersWasifur Rahman, Md. Kamrul Hasan, Sangwu Lee, AmirAli Bagher Zadeh et al.ACL 2020 · 584 citations
- Multi-level Cross-view Contrastive Learning for Knowledge-aware Recommender SystemDing Zou, Wei Wei, Xian-Ling Mao, Ziyang Wang et al.SIGIR 2022 · 226 citations
- Factorized Contrastive Learning: Going Beyond Multi-view RedundancyPaul Pu Liang, Zihao Deng, Martin Q. Ma, James Y. Zou et al.NeurIPS 2023 · 137 citations
- Sequential Latent Spaces for Modeling the Intention During Diverse Image CaptioningJyoti Aneja, Harsh Agrawal, Dhruv Batra, Alexander G. SchwingICCV 2019 · 71 citations
- Token-Level Contrastive Learning with Modality-Aware Prompting for Multimodal Intent RecognitionQianrui Zhou, Hua Xu, Hao Li, Hanlei Zhang et al.AAAI 2024 · 45 citations
Related papers
- CL-DMDF: Dynamic Multimodal Data Fusion Model Based on Contrastive LearningDong Li, Lingling Zhang, Binghao Han, Linlin Ding et al.AAAI 2026
- Audio-Visual Adaptive Fusion Network for Question Answering Based on Contrastive LearningXujian Zhao, Yixin Wang, Peiquan JinAAAI 2025 · 4 citations
- Cross-modality Representation Interactive Learning for Multimodal Sentiment AnalysisJian Huang, Yanli Ji, Yang Yang, Heng Tao ShenACM MM 2023 · 17 citations
- Contextual Augmented Global Contrast for Multimodal Intent RecognitionKaili Sun, Zhiwen Xie, Mang Ye, Huyin ZhangCVPR 2024 · 19 citations
- Adaptive Contrastive Learning on Multimodal Transformer for Review Helpfulness PredictionThong Nguyen, Xiaobao Wu, Anh Tuan Luu, Zhen Hai et al.EMNLP 2022 · 8 citations
