Multimodal Prompt Transformer with Hybrid Contrastive Learning for Emotion Recognition in Conversation
Shihao Zou, Xianying Huang, Xudong Shen
Abstract
Emotion Recognition in Conversation (ERC) plays an important role in driving the development of human-machine interaction. Emotions can exist in multiple modalities, and multimodal ERC mainly faces two problems: (1) the noise problem in the cross-modal information fusion process, and (2) the prediction problem of less sample emotion labels that are semantically similar but different categories. To address these issues and fully utilize the features of each modality, we adopted the following strategies: first, deep emotion cues extraction was performed on modalities with strong representation ability, and feature filters were designed as multimodal prompt information for modalities with weak representation ability. Then, we designed a Multimodal Prompt Transformer (MPT) to perform cross-modal information fusion. MPT embeds multimodal fusion information into each attention layer of the Transformer, allowing prompt information to participate in encoding textual features and being fused with multi-level textual information to obtain better multimodal fusion features. Finally, we used the Hybrid Contrastive Learning (HCL) strategy to optimize the model's ability to handle labels with few samples. This strategy uses unsupervised contrastive learning to improve the representation ability of multimodal fusion and supervised contrastive learning to mine the information of labels with few samples. Experimental results show that our proposed model outperforms state-of-the-art models in ERC on two benchmark datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- Revisiting Multimodal Emotion Recognition in Conversation from the Perspective of Graph SpectrumWei Ai, Fuchen Zhang, Yuntao Shou, Tao Meng et al.AAAI 2025 · 64 citations
- Unveiling the Cognitive Compass: Theory-of-Mind-Guided Multimodal Emotion ReasoningMeng Luo, Bobo Li, Shanqing Xu, Shize Zhang et al.ICLR 2026 · 10 citations
- Cross-modal Prompting for Balanced Incomplete Multi-modal Emotion RecognitionWenjue He, Xiaofeng Zhu, Zheng ZhangAAAI 2026 · 1 citation
- GS-Fuse: Granger-Supervised Gated Fusion and Multi-Granularity Alignment for Event-Driven Financial ForecastingYang Zhang, En Chun, Ziyun Mao, Yulu Wu et al.KDD 2026
- CoRiM: Conflict-driven Risk Minimization for Dynamic Multimodal FusionShihao Zou, Wei WeiCVPR 2026
Builds on13
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- DialogXL: All-in-One XLNet for Multi-Party Conversation Emotion RecognitionWeizhou Shen, Junqing Chen, Xiaojun Quan, Zhixian XieAAAI 2021 · 251 citations
- Relation-aware Graph Attention Networks with Relational Position Encodings for Emotion Recognition in ConversationsTaichi Ishiwatari, Yuki Yasuda, Taro Miyazaki, Jun GotoEMNLP 2020 · 201 citations
- Contrast and Generation Make BART a Good Dialogue Emotion RecognizerShimin Li, Hang Yan, Xipeng QiuAAAI 2022 · 121 citations
Related papers
- MultiEMO: An Attention-Based Correlation-Aware Multimodal Fusion Framework for Emotion Recognition in ConversationsTao Shi, Shao-Lun HuangACL 2023 · 76 citations
- Beyond Single Emotion: Multi-label Approach to Conversational Emotion RecognitionYujin Kang, Yoon-Sik ChoAAAI 2025 · 7 citations
- A Unimodal Valence-Arousal Driven Contrastive Learning Framework for Multimodal Multi-Label Emotion RecognitionWenjie Zheng, Jianfei Yu, Rui XiaACM MM 2024 · 8 citations
- UniMSE: Towards Unified Multimodal Sentiment Analysis and Emotion RecognitionGuimin Hu, Ting-En Lin, Yi Zhao, Guangming Lu et al.EMNLP 2022 · 206 citations
- A Cross-Modality Context Fusion and Semantic Refinement Network for Emotion Recognition in ConversationXiaoheng Zhang, Yang LiACL 2023 · 47 citations
