Revisiting Multimodal Emotion Recognition in Conversation from the Perspective of Graph Spectrum
Wei Ai, Fuchen Zhang, Yuntao Shou, Tao Meng, Haowen Chen, Keqin Li
Abstract
Efficiently capturing consistent and complementary semantic features in a multimodal conversation context is crucial for Multimodal Emotion Recognition in Conversation (MERC). Existing methods mainly use graph structures to model dialogue context semantic dependencies and employ Graph Neural Networks (GNN) to capture multimodal semantic features for emotion recognition. However, these methods are limited by some inherent characteristics of GNN, such as over-smoothing and low-pass filtering, resulting in the inability to learn long-distance consistency information and complementary information efficiently. Since consistency and complementarity information correspond to low-frequency and high-frequency information, respectively, this paper revisits the problem of multimodal emotion recognition in conversation from the perspective of the graph spectrum. Specifically, we propose a Graph-Spectrumbased Multimodal Consistency and Complementary collaborative learning framework GS-MCC. First, GS-MCC uses a sliding window to construct a multimodal interaction graph to model conversational relationships and uses efficient Fourier graph operators to extract long-distance high-frequency and low-frequency information, respectively. Then, GS-MCC uses contrastive learning to construct self-supervised signals that reflect complementarity and consistent semantic collaboration with high and low-frequency signals, thereby improving the ability of high and low-frequency information to reflect real emotions. Finally, GS-MCC inputs the collaborative high and low-frequency information into the MLP network and softmax function for emotion prediction. Extensive experiments have proven the superiority of the GS-MCC architecture proposed in this paper on two benchmark data sets. CCS CONCEPTS • Computing methodologies → Discourse, dialogue and pragmatics; Non-negative matrix factorization; • Theory of computation → Fixed parameter tractability.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 34e5f500-25d2-4c8f-b901-7497db5da1a0Cited by top-tier papers2
- Graph Domain Adaptation With Dual-Branch Encoder and Two-Level Alignment for Whole Slide Image-Based Survival PredictionYuntao Shou, Xiangyong Cao, Peiqiang Yan, Qiaohui et al.ICCV 2025 · 3 citations
- Beyond Missing Modalities: Hypergraph Conditioned Diffusion for Uncertainty-Aware Multimodal Emotion RecognitionXihang Qiu, Yuhao Fang, Qing Zhou, Bin Zhai et al.CVPR 2026
Builds on17
- FourierGNN: Rethinking Multivariate Time Series Forecasting from a Pure Graph PerspectiveKun Yi, Qi Zhang, Wei Fan, Hui He et al.NeurIPS 2023 · 359 citations
- Disentangled Representation Learning for Multimodal Emotion RecognitionDingkang Yang, Shuai Huang, Haopeng Kuang, Yangtao Du et al.ACM MM 2022 · 260 citations
- Relation-aware Graph Attention Networks with Relational Position Encodings for Emotion Recognition in ConversationsTaichi Ishiwatari, Yuki Yasuda, Taro Miyazaki, Jun GotoEMNLP 2020 · 201 citations
- Revisiting Graph Contrastive Learning from the Perspective of Graph SpectrumNian Liu, Xiao Wang, Deyu Bo, Chuan Shi et al.NeurIPS 2022 · 102 citations
- Distribution-Consistent Modal Recovering for Incomplete Multimodal LearningYuanzhi Wang, Zhen Cui, Yong LiICCV 2023 · 101 citations
Related papers
- Multivariate, Multi-Frequency and Multimodal: Rethinking Graph Neural Networks for Emotion Recognition in ConversationFeiyu Chen, Jie Shao, Shuyuan Zhu, Heng Tao ShenCVPR 2023
- Joyful: Joint Modality Fusion and Graph Contrastive Learning for Multimoda Emotion RecognitionDongyuan Li, Yusong Wang, Kotaro Funakoshi, Manabu OkumuraEMNLP 2023 · 46 citations
- Multimodal Fusion via Hypergraph Autoencoder and Contrastive Learning for Emotion Recognition in ConversationZijian Yi, Ziming Zhao, Zhishu Shen, Tiehua ZhangACM MM 2024 · 32 citations
- A Cross-Modality Context Fusion and Semantic Refinement Network for Emotion Recognition in ConversationXiaoheng Zhang, Yang LiACL 2023 · 47 citations
- MMGCN: Multimodal Fusion via Deep Graph Convolution Network for Emotion Recognition in ConversationJingwen Hu, Yuchen Liu, Jinming Zhao, Qin JinACL 2021
