Dual-Path Dynamic Fusion with Learnable Query for Multimodal Sentiment Analysis
Miao Zhou, Lina Yang, Thomas Wu, Dongnan Yang, Xinru Zhang
Abstract
Multimodal Sentiment Analysis (MSA) is the task of understanding human emotions by analyzing a combination of different data sources, such as text, audio, and visual inputs. Although recent advances have improved emotion modeling across modalities, existing methods still struggle with two fundamental challenges: balancing global and fine-grained sentiment contributions, and over-reliance on the text modality. To address these issues, we propose DPDF-LQ (Dual-Path Dynamic Fusion with Learnable Query), an architecture that processes inputs through two complementary paths: global and local. The global path is responsible for establishing cross-modal dependencies, while the local path captures finegrained representations. Additionally, we introduce the key module Dynamic Global Learnable Query Attention (DGLQA) in the global path, which dynamically allocates weights to each modality to capture their relevant features and learn global representations. Extensive experiments on the CMU-MOSI and CMU-MOSEI benchmarks demonstrate that DPDF-LQ achieves state-of-the-art performance, particularly in fine-grained sentiment prediction by effectively combining global and local features. Our code will be released at https: //github.com/ZhouMiaoGX/DPDF-LQ .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b947ca0f-861c-47a7-8f33-0143a6e260d0Builds on9
- MISA: Modality-Invariant and -Specific Representations for Multimodal Sentiment AnalysisDevamanyu Hazarika, Roger Zimmermann, Soujanya PoriaACM MM 2020 · 1,037 citations
- Learning Modality-Specific Representations with Self-Supervised Multi-Task Learning for Multimodal Sentiment AnalysisWenmeng Yu, Hua Xu, Ziqi Yuan, Jiele WuAAAI 2021 · 737 citations
- CH-SIMS: A Chinese Multimodal Sentiment Analysis Dataset with Fine-grained Annotation of ModalityWenmeng Yu, Hua Xu, Fanyang Meng, Yilin Zhu et al.ACL 2020 · 376 citations
- Disentangled Representation Learning for Multimodal Emotion RecognitionDingkang Yang, Shuai Huang, Haopeng Kuang, Yangtao Du et al.ACM MM 2022 · 260 citations
- Transformer-based Feature Reconstruction Network for Robust Multimodal Sentiment AnalysisZiqi Yuan, Wei Li, Hua Xu, Wenmeng YuACM MM 2021 · 186 citations
Related papers
- DDSE: A Decoupled Dual-Stream Enhanced Framework for Multimodal Sentiment Analysis with Text-Centric SSMShenjie Jiang, Zhuoyu Wang, Xuecheng Wu, Hongru Ji et al.ACM MM 2025 · 4 citations
- DLF: Disentangled-Language-Focused Multimodal Sentiment AnalysisPan Wang, Qiang Zhou, Yawen Wu, Tianlong Chen et al.AAAI 2025 · 84 citations
- GLoMo: Global-Local Modal Fusion for Multimodal Sentiment AnalysisYan Zhuang, Yanru Zhang, Zheng Hu, Xiaoyue Zhang et al.ACM MM 2024 · 26 citations
- PSA-MF: Personality-Sentiment Aligned Multi-Level Fusion for Multimodal Sentiment AnalysisHeng Xie, Kang Zhu, Zhengqi Wen, Jianhua Tao et al.AAAI 2026 · 1 citation
- DiffuFuse: Diffusion-Driven Dual-Stream Fusion Framework for Multimodal Sentiment AnalysisXiongjian Lv, Yimin Wen, Hang YuACM MM 2025
