Lune

ACM MM2025Top-tier venue

DiffuFuse: Diffusion-Driven Dual-Stream Fusion Framework for Multimodal Sentiment Analysis

Xiongjian Lv, Yimin Wen, Hang Yu

2025Year
1Top-tier citations

Abstract

Multimodal Sentiment Analysis (MSA) aims to integrate textual, audio, and visual data to capture nuanced sentimental cues. Although text dominates in existing approaches, audio and visual modalities inherently contain both shared semantics (overlapping with text) and private semantics. Existing methods struggle to precisely find semantic boundaries and lack explicit mechanisms for modeling interaction between shared/private semantics and different modalities. To address this, we propose DiffuFuse, a framework that uses a diffusion denoising model to leverage textual information to predict shared semantic features, dynamically and adaptively delineate semantic boundaries for non-textual features, and employs a dual-stream fusion strategy to accurately model the interactions between different modalities and semantic types. Finally, adopt an orthogonal projection method to reduce redundancy and eliminate overlapping information between the two streams. DiffuFuse is evaluated on the MOSI and MOSEI datasets, and the experimental results demonstrate that our proposed DiffuFuse achieves superior performance.

Ask about this paper

Ask your agent about it.

Lune has read the top-tier papers around this one, so every answer names the papers it rests on.

Questions to start from

Your agent calls

Lunesearch_papers

Ask in Lune

Free to start. No credit card required.

lune papers get 944f15ac-0198-44be-9b40-2d8ed205a088

Cited by top-tier papers1

Ask how each one uses it

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines