MultiEMO: An Attention-Based Correlation-Aware Multimodal Fusion Framework for Emotion Recognition in Conversations
Tao Shi, Shao-Lun Huang
Abstract
Emotion Recognition in Conversations (ERC) is an increasingly popular task in the Natural Language Processing community, which seeks to achieve accurate emotion classifications of utterances expressed by speakers during a conversation. Most existing approaches focus on modeling speaker and contextual information based on the textual modality, while the complementarity of multimodal information has not been well leveraged, few current methods have sufficiently captured the complex correlations and mapping relationships across different modalities. Furthermore, existing state-ofthe-art ERC models have difficulty classifying minority and semantically similar emotion categories. To address these challenges, we propose a novel attention-based correlation-aware multimodal fusion framework named MultiEMO, which effectively integrates multimodal cues by capturing cross-modal mapping relationships across textual, audio and visual modalities based on bidirectional multi-head crossattention layers. The difficulty of recognizing minority and semantically hard-to-distinguish emotion classes is alleviated by our proposed Sample-Weighted Focal Contrastive (SWFC) loss. Extensive experiments on two benchmark ERC datasets demonstrate that our MultiEMO framework consistently outperforms existing state-of-the-art approaches in all emotion categories on both datasets, the improvements in minority and semantically similar emotions are especially significant. * Corresponding author. potentials in social media analysis (Chatterjee et al., 2019) , health care services (Hu et al., 2021b), empathetic systems (Jiao et al., 2020) , and so on.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8909ed81-eda9-4512-a1d4-c367062e1019Cited by top-tier papers11
- Conversation Understanding using Relational Temporal Graph Neural Networks with Auxiliary Cross-Modality InteractionCam-Van Thi Nguyen, Anh-Tuan Mai, The-Son Le, Hai-Dang Kieu et al.EMNLP 2023 · 34 citations
- PanoSent: A Panoptic Sextuple Extraction Benchmark for Multimodal Conversational Aspect-based Sentiment AnalysisMeng Luo, Hao Fei, Bobo Li, Shengqiong Wu et al.ACM MM 2024 · 23 citations
- Text-Guided Fine-grained Counterfactual Inference for Short Video Fake News DetectionLinlin Zong, Wenmin Lin, Jiahui Zhou, Xinyue Liu et al.AAAI 2025 · 6 citations
- Grounding Emotion Recognition with Visual Prototypes: VEGA - Revisiting CLIP in MERCGuanyu Hu, Dimitrios Kollias, Xinyu YangACM MM 2025 · 5 citations
- ECERC: Evidence-Cause Attention Network for Multi-Modal Emotion Recognition in ConversationTao Zhang, Zhenhua TanACL 2025 · 4 citations
Builds on8
- UniMSE: Towards Unified Multimodal Sentiment Analysis and Emotion RecognitionGuimin Hu, Ting-En Lin, Yi Zhao, Guangming Lu et al.EMNLP 2022 · 206 citations
- Real-Time Emotion Recognition via Attention Gated Hierarchical Memory NetworkWenxiang Jiao, Michael R. Lyu, Irwin KingAAAI 2020 · 149 citations
- Unleashing the Power of Contrastive Self-Supervised Visual Models via Contrast-Regularized Fine-TuningYifan Zhang, Bryan Hooi, Dapeng Hu, Jian Liang et al.NeurIPS 2021 · 82 citations
- Quantum-inspired Neural Network for Conversational Emotion RecognitionQiuchi Li, Dimitris Gkoumas, Alessandro Sordoni, Jian-Yun Nie et al.AAAI 2021 · 59 citations
- Transformer Feed-Forward Layers Are Key-Value MemoriesMor Geva, Roei Schuster, Jonathan Berant, Omer LevyEMNLP 2021 · 33 citations
Related papers
- Multimodal Prompt Transformer with Hybrid Contrastive Learning for Emotion Recognition in ConversationShihao Zou, Xianying Huang, Xudong ShenACM MM 2023 · 24 citations
- Disentangled Representation Learning for Multimodal Emotion RecognitionDingkang Yang, Shuai Huang, Haopeng Kuang, Yangtao Du et al.ACM MM 2022 · 260 citations
- Emotion-Wheel-Guided Audio-Referred Text Representation for Multimodal Emotion Recognition in ConversationEunseon Seong, Harim Lee, Dahye Kim, Changhyun Kim et al.ACL 2026
- Beyond Single Emotion: Multi-label Approach to Conversational Emotion RecognitionYujin Kang, Yoon-Sik ChoAAAI 2025 · 7 citations
- VAEmo: Efficient Representation Learning for Visual-Audio Emotion With Knowledge InjectionHao Cheng, Zhiwei Zhao, Yichao He, Zhenzhen Hu et al.ACM MM 2025 · 9 citations
