Pairwise Emotional Relationship Recognition in Drama Videos: Dataset and Benchmark
Xun Gao, Yin Zhao, Jie Zhang, Longjun Cai
摘要
Recognizing the emotional state of people is a basic but challenging task in video understanding. In this paper, we propose a new task in this field, named Pairwise Emotional Relationship Recognition (PERR). This task aims to recognize the emotional relationship between the two interactive characters in a given video clip. It is different from the traditional emotion and social relation recognition task. Varieties of information, consisting of character appearance, behaviors, facial emotions, dialogues, background music as well as subtitles contribute differently to the final results, which makes the task more challenging but meaningful in developing more advanced multi-modal models. To facilitate the task, we develop a new dataset called Emotional RelAtionship of inTeractiOn (ERATO) based on dramas and movies. ERATO is a large-scale multi-modal dataset for PERR task, which has 31,182 video clips, lasting about 203 video hours. Different from the existing datasets, ERATO contains interaction-centric videos with multi-shots, varied video length, and multiple modalities including visual, audio and text. As a minor contribution, we propose a baseline model composed of Synchronous Modal-Temporal Attention (SMTA) unit to fuse the multi-modal information for the PERR task. In contrast to other prevailing attention mechanisms, our proposed SMTA can steadily improve the performance by about 1%. We expect the ER-ATO as well as our proposed SMTA to open up a new way for PERR task in video understanding and further improve the research of multi-modal fusion methodology.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- VidEmo: Affective-Tree Reasoning for Emotion-Centric Video Foundation ModelsZhicheng Zhang, Weicheng Wang, Yongjie Zhu, Wenyu Qin 等NeurIPS 2025 · 被引用 11 次
- MART: Masked Affective RepresenTation Learning via Masked Temporal Distribution DistillationZhicheng Zhang, Pancheng Zhao, Eunil Park, Jufeng YangCVPR 2024 · 被引用 11 次
它引用的顶会 Paper6
- VideoBERT: A Joint Model for Video and Language Representation LearningChen Sun, Austin Myers, Carl Vondrick, Kevin Murphy 等ICCV 2019 · 被引用 1,396 次
- Context-Aware Emotion Recognition NetworksJiyoung Lee, Seungryong Kim, Sunok Kim, Jungin Park 等ICCV 2019 · 被引用 285 次
- Parameter Efficient Multimodal Transformers for Video Representation LearningSangho Lee, Youngjae Yu, Gunhee Kim, Thomas M. Breuel 等ICLR 2021 · 被引用 90 次
- EmotiCon: Context-Aware Multimodal Emotion Recognition Using Frege's PrincipleTrisha Mittal, Pooja Guhan, Uttaran Bhattacharya, Rohan Chandra 等CVPR 2020
- Learning Interactions and Relationships Between Movie CharactersAnna Kukleva, Makarand Tapaswi, Ivan LaptevCVPR 2020
相关 Paper
- MEmoR: A Dataset for Multimodal Emotion Reasoning in VideosGuangyao Shen, Xin Wang, Xuguang Duan, Hongzhi Li 等ACM MM 2020 · 被引用 38 次
- Towards Emotion-aided Multi-modal Dialogue Act ClassificationTulika Saha, Aditya Prakash Patra, Sriparna Saha, Pushpak BhattacharyyaACL 2020 · 被引用 63 次
- MEDIC: A Multimodal Empathy Dataset in CounselingZhouan Zhu, Chenguang Li, Jicai Pan, Xin Li 等ACM MM 2023 · 被引用 9 次
- MultiEMO: An Attention-Based Correlation-Aware Multimodal Fusion Framework for Emotion Recognition in ConversationsTao Shi, Shao-Lun HuangACL 2023 · 被引用 76 次
- Violin: A Large-Scale Dataset for Video-and-Language InferenceJingzhou Liu, Wenhu Chen, Yu Cheng, Zhe Gan 等CVPR 2020
