Lune

ACM MM2025顶会

DiffuFuse: Diffusion-Driven Dual-Stream Fusion Framework for Multimodal Sentiment Analysis

Xiongjian Lv, Yimin Wen, Hang Yu

2025年份
1顶会引用

摘要

Multimodal Sentiment Analysis (MSA) aims to integrate textual, audio, and visual data to capture nuanced sentimental cues. Although text dominates in existing approaches, audio and visual modalities inherently contain both shared semantics (overlapping with text) and private semantics. Existing methods struggle to precisely find semantic boundaries and lack explicit mechanisms for modeling interaction between shared/private semantics and different modalities. To address this, we propose DiffuFuse, a framework that uses a diffusion denoising model to leverage textual information to predict shared semantic features, dynamically and adaptively delineate semantic boundaries for non-textual features, and employs a dual-stream fusion strategy to accurately model the interactions between different modalities and semantic types. Finally, adopt an orthogonal projection method to reduce redundancy and eliminate overlapping information between the two streams. DiffuFuse is evaluated on the MOSI and MOSEI datasets, and the experimental results demonstrate that our proposed DiffuFuse achieves superior performance.

问问这篇 Paper

问问你的智能体。

Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。

可以从这些问题问起

智能体调用

Lunesearch_papers

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper1

问问它们各自怎么用它

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖