Multimodal Fusion via Hypergraph Autoencoder and Contrastive Learning for Emotion Recognition in Conversation
Zijian Yi, Ziming Zhao, Zhishu Shen, Tiehua Zhang
Abstract
Multimodal emotion recognition in conversation (MERC) seeks to identify the speakers' emotions expressed in each utterance, offering significant potential across diverse fields. The challenge of MERC lies in balancing speaker modeling and context modeling, encompassing both long-distance and short-distance contexts, as well as addressing the complexity of multimodal information fusion. Recent research adopts graph-based methods to model intricate conversational relationships effectively. Nevertheless, the majority of these methods utilize a fixed fully connected structure to link all utterances, relying on convolution to interpret complex context. This approach can inherently heighten the redundancy in contextual messages and excessive graph network smoothing, particularly in the context of long-distance conversations. To address this issue, we propose a framework that dynamically adjusts hypergraph connections by variational hypergraph autoencoder (VHGAE), and employs contrastive learning to mitigate uncertainty factors during the reconstruction process. Experimental results demonstrate the effectiveness of our proposal against the state-of-the-art methods on IEMOCAP and MELD datasets. We release the code to support the reproducibility of this work at https://github.com/yzjred/-HAUCL.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 23114d75-3f30-43a4-9e94-19fb90b08022Cited by top-tier papers4
- ECERC: Evidence-Cause Attention Network for Multi-Modal Emotion Recognition in ConversationTao Zhang, Zhenhua TanACL 2025 · 4 citations
- Hypergraph-based Zero-shot Multi-modal Product Attribute Value ExtractionJiazhen Hu, Jiaying Gong, Hongda Shen, Hoda EldardiryWWW 2025 · 4 citations
- Cross-Space Synergy: A Unified Framework for Multimodal Emotion Recognition in ConversationXiaosen Lyu, Jiayu Xiong, Yuren Chen, Wanlong Wang et al.AAAI 2026
- Personality-guided Public-Private Domain Disentangled Hypergraph-Former Network for Multimodal Depression DetectionChangzeng Fu, Shiwen Zhao, Yunze Zhang, Zhongquan Jian et al.AAAI 2026
Builds on9
- DropEdge: Towards Deep Graph Convolutional Networks on Node ClassificationYu Rong, Wenbing Huang, Tingyang Xu, Junzhou HuangICLR 2020 · 1,599 citations
- Disentangled Representation Learning for Multimodal Emotion RecognitionDingkang Yang, Shuai Huang, Haopeng Kuang, Yangtao Du et al.ACM MM 2022 · 260 citations
- Self-Supervised Hypergraph Transformer for Recommender SystemsLianghao Xia, Chao Huang, Chuxu ZhangKDD 2022 · 142 citations
- Augmentations in Hypergraph Contrastive Learning: Fabricated and GenerativeTianxin Wei, Yuning You, Tianlong Chen, Yang Shen et al.NeurIPS 2022 · 96 citations
- Unlocking the Power of Multimodal Learning for Emotion Recognition in ConversationYunxiao Wang, Meng Liu, Zhe Li, Yupeng Hu et al.ACM MM 2023 · 16 citations
Related papers
- MMGCN: Multimodal Fusion via Deep Graph Convolution Network for Emotion Recognition in ConversationJingwen Hu, Yuchen Liu, Jinming Zhao, Qin JinACL 2021
- Multimodal Emotion Recognition Calibration in ConversationsGeng Tu, Feng Xiong, Bin Liang, Hui Wang et al.ACM MM 2024 · 12 citations
- Joyful: Joint Modality Fusion and Graph Contrastive Learning for Multimoda Emotion RecognitionDongyuan Li, Yusong Wang, Kotaro Funakoshi, Manabu OkumuraEMNLP 2023 · 46 citations
- BIG-FUSION: Brain-Inspired Global-Local Context Fusion Framework for Multimodal Emotion Recognition in ConversationsYusong Wang, Xuanye Fang, Huifeng Yin, Dongyuan Li et al.AAAI 2025 · 11 citations
- Dynamic Interactive Bimodal Hypergraph Networks for Emotion Recognition in ConversationsXuping Chen, Wuzhen ShiAAAI 2025 · 3 citations
