Transformer-based Label Set Generation for Multi-modal Multi-label Emotion Detection
Xincheng Ju, Dong Zhang, Junhui Li, Guodong Zhou
摘要
Multi-modal utterance-level emotion detection has been a hot research topic in both multi-modal analysis and natural language processing communities. Different from traditional single-label multi-modal sentiment analysis, typical multi-modal emotion detection is naturally a multi-label problem where an utterance often contains multiple emotions. Existing studies normally focus on multi-modal fusion only and transform multi-label emotion classification into multiple binary classification problem independently. As a result, existing studies largely ignore two kinds of important dependency information: (1) Modality-to-label dependency, where different emotions can be inferred from different modalities, that is, different modalities contribute differently to each potential emotion. (2) Label-to-label dependency, where some emotions are more likely to coexist than those conflicting emotions. To simultaneously model above two kinds of dependency, we propose a unified approach, namely multi-modal emotion set generation network (MESGN) to generate an emotion set for an utterance. Specifically, we first employ a cross-modal transformer encoder to capture cross-modal interactions among different modalities, and a standard transformer encoder to capture temporal information for each modality-specific sequence given previous interactions. Then, we design a transformer-based discriminative decoding module equipped with modality-to-label attention to handle the modality-to-label dependency. In the meanwhile, we employ a reinforced decoding algorithm with self-critic learning to handle the label-to-label dependency. Finally, we validate the proposed MESGN architecture on a word-level aligned and unaligned multi-modal dataset. Detailed experimentation shows that our proposed MESGN architecture can effectively improve the performance of multi-modal multi-label emotion detection.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper12
- Multi-modal Graph Fusion for Named Entity Recognition with Targeted Visual GuidanceDong Zhang, Suzhong Wei, Shoushan Li, Hanqian Wu 等AAAI 2021 · 被引用 240 次
- Joint Multi-modal Aspect-Sentiment Analysis with Auxiliary Cross-modal Relation DetectionXincheng Ju, Dong Zhang, Rong Xiao, Junhui Li 等EMNLP 2021 · 被引用 130 次
- Tailor Versatile Multi-Modal Learning for Multi-Label Emotion RecognitionYi Zhang, Mingyuan Chen, Jundong Shen, Chongjun WangAAAI 2022 · 被引用 92 次
- Joint Multimodal Entity-Relation Extraction Based on Edge-Enhanced Graph Alignment Network and Word-Pair Relation TaggingLi Yuan, Yi Cai, Jin Wang, Qing LiAAAI 2023 · 被引用 89 次
- DocTr: Document Image Transformer for Geometric Unwarping and Illumination CorrectionHao Feng, Yuechen Wang, Wengang Zhou, Jiajun Deng 等ACM MM 2021 · 被引用 66 次
相关 Paper
- Multi-modal Multi-label Emotion Detection with Modality and Label DependenceDong Zhang, Xincheng Ju, Junhui Li, Shoushan Li 等EMNLP 2020 · 被引用 45 次
- Multi-modal Multi-label Emotion Recognition with Heterogeneous Hierarchical Message PassingDong Zhang, Xincheng Ju, Wei Zhang, Junhui Li 等AAAI 2021 · 被引用 56 次
- Layer-wise Fusion with Modality Independence Modeling for Multi-modal Emotion RecognitionJun Sun, Shoukang Han, Yu-Ping Ruan, Xiaoning Zhang 等ACL 2023 · 被引用 23 次
- Conversation Understanding using Relational Temporal Graph Neural Networks with Auxiliary Cross-Modality InteractionCam-Van Thi Nguyen, Anh-Tuan Mai, The-Son Le, Hai-Dang Kieu 等EMNLP 2023 · 被引用 34 次
- Dynamically Adjust Word Representations Using Unaligned Multimodal InformationJiwei Guo, Jiajia Tang, Weichen Dai, Yu Ding 等ACM MM 2022 · 被引用 59 次
