Adaptive Multimodal Fusion: Dynamic Attention Allocation for Intent Recognition
Bo Hu, Kai Zhang, Yanghai Zhang, Yuyang Ye
摘要
In recent years, deep multimodal learning has seen significant advancements. However, there remains a lack of multimodal fusion methods capable of dynamically adjusting the weighting of information both within and across modalities based on input samples. In the domain of multimodal intent recognition, the text modality often contains the most relevant information for intent detection, while the audio and visual modalities provide comparatively less critical information. There is a significant variation in the density of important information across different modalities and samples. To address this challenge, we propose a Dynamic Attention Allocation Fusion (DAF) method with an adaptive network structure that dynamically allocates attention both within individual modalities and across multiple modalities. This approach enables the model to focus more effectively on the most informative modalities and their respective internal features. Furthermore, we introduce a multi-view contrastive learning framework based on DAF (MVCL-DAF). This framework uses distinct and isolated modules to process information from various modalities, taking inspiration from the way the human brain processes multimodal information. Each modality independently infers intent using its respective module, while DAF integrates the multimodal information to produce a comprehensive global intent prediction. The text modality, functioning as the primary modality due to its rich semantic content, guides the other modules in the multi-view contrastive learning process. Extensive experiments demonstrate that our approach significantly outperforms existing state-of-the-art methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Evolutionary Multimodal Reasoning via Hierarchical Semantic Representation for Intent RecognitionQianrui Zhou, Hua Xu, Yunjin Gu, Yifan Wang 等CVPR 2026 · 被引用 3 次
- Who Should I Trust? Explicit Confidence-Focused Multimodal Intent RecognitionYi Liu, Qimeng Yang, Lanlan LuAAAI 2026
- Inconsistency-aware Multimodal Schrödinger Bridge for Deepfake LocalizationJiayu Xiong, Jing Wang, Qi Zhang, Wanlong Wang 等CVPR 2026
- SeD-UD: An Influence-Driven and Hierarchically-Decoupled Information Bottleneck for Multimodal Intent RecognitionQin Li, Wenbo Zhang, Limei Liu, Han Peng 等CVPR 2026
它引用的顶会 Paper9
- Integrating Multimodal Information in Large Pretrained TransformersWasifur Rahman, Md. Kamrul Hasan, Sangwu Lee, AmirAli Bagher Zadeh 等ACL 2020 · 被引用 584 次
- Multi-level Cross-view Contrastive Learning for Knowledge-aware Recommender SystemDing Zou, Wei Wei, Xian-Ling Mao, Ziyang Wang 等SIGIR 2022 · 被引用 226 次
- Factorized Contrastive Learning: Going Beyond Multi-view RedundancyPaul Pu Liang, Zihao Deng, Martin Q. Ma, James Y. Zou 等NeurIPS 2023 · 被引用 137 次
- Sequential Latent Spaces for Modeling the Intention During Diverse Image CaptioningJyoti Aneja, Harsh Agrawal, Dhruv Batra, Alexander G. SchwingICCV 2019 · 被引用 71 次
- Token-Level Contrastive Learning with Modality-Aware Prompting for Multimodal Intent RecognitionQianrui Zhou, Hua Xu, Hao Li, Hanlei Zhang 等AAAI 2024 · 被引用 45 次
相关 Paper
- CL-DMDF: Dynamic Multimodal Data Fusion Model Based on Contrastive LearningDong Li, Lingling Zhang, Binghao Han, Linlin Ding 等AAAI 2026
- Audio-Visual Adaptive Fusion Network for Question Answering Based on Contrastive LearningXujian Zhao, Yixin Wang, Peiquan JinAAAI 2025 · 被引用 4 次
- Cross-modality Representation Interactive Learning for Multimodal Sentiment AnalysisJian Huang, Yanli Ji, Yang Yang, Heng Tao ShenACM MM 2023 · 被引用 17 次
- Contextual Augmented Global Contrast for Multimodal Intent RecognitionKaili Sun, Zhiwen Xie, Mang Ye, Huyin ZhangCVPR 2024 · 被引用 19 次
- Adaptive Contrastive Learning on Multimodal Transformer for Review Helpfulness PredictionThong Nguyen, Xiaobao Wu, Anh Tuan Luu, Zhen Hai 等EMNLP 2022 · 被引用 8 次
