MultiEMO: An Attention-Based Correlation-Aware Multimodal Fusion Framework for Emotion Recognition in Conversations
Tao Shi, Shao-Lun Huang
摘要
Emotion Recognition in Conversations (ERC) is an increasingly popular task in the Natural Language Processing community, which seeks to achieve accurate emotion classifications of utterances expressed by speakers during a conversation. Most existing approaches focus on modeling speaker and contextual information based on the textual modality, while the complementarity of multimodal information has not been well leveraged, few current methods have sufficiently captured the complex correlations and mapping relationships across different modalities. Furthermore, existing state-ofthe-art ERC models have difficulty classifying minority and semantically similar emotion categories. To address these challenges, we propose a novel attention-based correlation-aware multimodal fusion framework named MultiEMO, which effectively integrates multimodal cues by capturing cross-modal mapping relationships across textual, audio and visual modalities based on bidirectional multi-head crossattention layers. The difficulty of recognizing minority and semantically hard-to-distinguish emotion classes is alleviated by our proposed Sample-Weighted Focal Contrastive (SWFC) loss. Extensive experiments on two benchmark ERC datasets demonstrate that our MultiEMO framework consistently outperforms existing state-of-the-art approaches in all emotion categories on both datasets, the improvements in minority and semantically similar emotions are especially significant. * Corresponding author. potentials in social media analysis (Chatterjee et al., 2019) , health care services (Hu et al., 2021b), empathetic systems (Jiao et al., 2020) , and so on.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Conversation Understanding using Relational Temporal Graph Neural Networks with Auxiliary Cross-Modality InteractionCam-Van Thi Nguyen, Anh-Tuan Mai, The-Son Le, Hai-Dang Kieu 等EMNLP 2023 · 被引用 34 次
- PanoSent: A Panoptic Sextuple Extraction Benchmark for Multimodal Conversational Aspect-based Sentiment AnalysisMeng Luo, Hao Fei, Bobo Li, Shengqiong Wu 等ACM MM 2024 · 被引用 23 次
- Text-Guided Fine-grained Counterfactual Inference for Short Video Fake News DetectionLinlin Zong, Wenmin Lin, Jiahui Zhou, Xinyue Liu 等AAAI 2025 · 被引用 6 次
- Grounding Emotion Recognition with Visual Prototypes: VEGA - Revisiting CLIP in MERCGuanyu Hu, Dimitrios Kollias, Xinyu YangACM MM 2025 · 被引用 5 次
- ECERC: Evidence-Cause Attention Network for Multi-Modal Emotion Recognition in ConversationTao Zhang, Zhenhua TanACL 2025 · 被引用 4 次
它引用的顶会 Paper8
- UniMSE: Towards Unified Multimodal Sentiment Analysis and Emotion RecognitionGuimin Hu, Ting-En Lin, Yi Zhao, Guangming Lu 等EMNLP 2022 · 被引用 206 次
- Real-Time Emotion Recognition via Attention Gated Hierarchical Memory NetworkWenxiang Jiao, Michael R. Lyu, Irwin KingAAAI 2020 · 被引用 149 次
- Unleashing the Power of Contrastive Self-Supervised Visual Models via Contrast-Regularized Fine-TuningYifan Zhang, Bryan Hooi, Dapeng Hu, Jian Liang 等NeurIPS 2021 · 被引用 82 次
- Quantum-inspired Neural Network for Conversational Emotion RecognitionQiuchi Li, Dimitris Gkoumas, Alessandro Sordoni, Jian-Yun Nie 等AAAI 2021 · 被引用 59 次
- Transformer Feed-Forward Layers Are Key-Value MemoriesMor Geva, Roei Schuster, Jonathan Berant, Omer LevyEMNLP 2021 · 被引用 33 次
相关 Paper
- Multimodal Prompt Transformer with Hybrid Contrastive Learning for Emotion Recognition in ConversationShihao Zou, Xianying Huang, Xudong ShenACM MM 2023 · 被引用 24 次
- Disentangled Representation Learning for Multimodal Emotion RecognitionDingkang Yang, Shuai Huang, Haopeng Kuang, Yangtao Du 等ACM MM 2022 · 被引用 260 次
- Emotion-Wheel-Guided Audio-Referred Text Representation for Multimodal Emotion Recognition in ConversationEunseon Seong, Harim Lee, Dahye Kim, Changhyun Kim 等ACL 2026
- Beyond Single Emotion: Multi-label Approach to Conversational Emotion RecognitionYujin Kang, Yoon-Sik ChoAAAI 2025 · 被引用 7 次
- VAEmo: Efficient Representation Learning for Visual-Audio Emotion With Knowledge InjectionHao Cheng, Zhiwei Zhao, Yichao He, Zhenzhen Hu 等ACM MM 2025 · 被引用 9 次
