A Facial Expression-Aware Multimodal Multi-task Learning Framework for Emotion Recognition in Multi-party Conversations
Wenjie Zheng, Jianfei Yu, Rui Xia, Shijin Wang
摘要
Multimodal Emotion Recognition in Multiparty Conversations (MERMC) has recently attracted considerable attention. Due to the complexity of visual scenes in multi-party conversations, most previous MERMC studies mainly focus on text and audio modalities while ignoring visual information. Recently, several works proposed to extract face sequences as visual features and have shown the importance of visual information in MERMC. However, given an utterance, the face sequence extracted by previous methods may contain multiple people's faces, which will inevitably introduce noise to the emotion prediction of the real speaker. To tackle this issue, we propose a two-stage framework named Facial expressionaware Multimodal Multi-Task learning (Fa-cialMMT). Specifically, a pipeline method is first designed to extract the face sequence of the real speaker of each utterance, which consists of multimodal face recognition, unsupervised face clustering, and face matching. With the extracted face sequences, we propose a multimodal facial expression-aware emotion recognition model, which leverages the frame-level facial emotion distributions to help improve utterance-level emotion recognition based on multi-task learning. Experiments demonstrate the effectiveness of the proposed FacialMMT framework on the benchmark MELD dataset. The source code is publicly released at https: //github.com/NUSTM/FacialMMT .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- PanoSent: A Panoptic Sextuple Extraction Benchmark for Multimodal Conversational Aspect-based Sentiment AnalysisMeng Luo, Hao Fei, Bobo Li, Shengqiong Wu 等ACM MM 2024 · 被引用 23 次
- Grounding Emotion Recognition with Visual Prototypes: VEGA - Revisiting CLIP in MERCGuanyu Hu, Dimitrios Kollias, Xinyu YangACM MM 2025 · 被引用 5 次
- Knowledge-Aligned Counterfactual-Enhancement Diffusion Perception for Unsupervised Cross-Domain Visual Emotion RecognitionWen Yin, Yong Wang, Guiduo Duan, Dongyang Zhang 等CVPR 2025
- D2SP: Dynamic Dual-Stage Purification Framework for Dual Noise Mitigation in Vision-based Affective RecognitionHaoran Wang, Xinji Mai, Zeng Tao, Xuan Tong 等CVPR 2025
它引用的顶会 Paper12
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 被引用 9,451 次
- DFEW: A Large-Scale Database for Recognizing Dynamic Facial Expressions in the WildXingxun Jiang, Yuan Zong, Wenming Zheng, Chuangao Tang 等ACM MM 2020 · 被引用 205 次
- Is Someone Speaking?: Exploring Long-term Temporal Features for Audio-visual Active Speaker DetectionRuijie Tao, Zexu Pan, Rohan Kumar Das, Xinyuan Qian 等ACM MM 2021 · 被引用 154 次
- FERV39k: A Large-Scale Multi-Scene Dataset for Facial Expression Recognition in VideosYan Wang, Yixuan Sun, Yiwen Huang, Zhongying Liu 等CVPR 2022 · 被引用 107 次
相关 Paper
- Friends-MMC: A Dataset for Multi-modal Multi-party Conversation UnderstandingYueqian Wang, Xiaojun Meng, Yuxuan Wang, Jianxin Liang 等AAAI 2025 · 被引用 6 次
- MultiEMO: An Attention-Based Correlation-Aware Multimodal Fusion Framework for Emotion Recognition in ConversationsTao Shi, Shao-Lun HuangACL 2023 · 被引用 76 次
- ECERC: Evidence-Cause Attention Network for Multi-Modal Emotion Recognition in ConversationTao Zhang, Zhenhua TanACL 2025 · 被引用 4 次
- Towards Emotion-aided Multi-modal Dialogue Act ClassificationTulika Saha, Aditya Prakash Patra, Sriparna Saha, Pushpak BhattacharyyaACL 2020 · 被引用 63 次
- Conversation Understanding using Relational Temporal Graph Neural Networks with Auxiliary Cross-Modality InteractionCam-Van Thi Nguyen, Anh-Tuan Mai, The-Son Le, Hai-Dang Kieu 等EMNLP 2023 · 被引用 34 次
