MEmoR: A Dataset for Multimodal Emotion Reasoning in Videos
Guangyao Shen, Xin Wang, Xuguang Duan, Hongzhi Li, Wenwu Zhu
摘要
Humans can perceive subtle emotions from various cues and contexts, even without hearing or seeing others. However, existing video datasets mainly focus on recognizing the emotions of the speakers from complete modalities. In this work, we present the task of multimodal emotion reasoning in videos. Beyond directly recognizing emotions from multimodal signals of target persons, this task requires a machine capable of reasoning about human emotions from the contexts and surrounding world. To facilitate the study towards this task, we introduce a new dataset, MEmoR, that provides fine-grained emotion annotations for both speakers and non-speakers. The videos in MEmoR are collected from TV shows closely in real-life scenarios. In these videos, while speakers may be non-visually described, non-speakers always deliver no audio-textual signals and are often visually inconspicuous. This modality-missing characteristic makes MEmoR a more practical yet challenging testbed for multimodal emotion reasoning. In support of various reasoning behaviors, the proposed MEmoR dataset provides both short-term contexts and external knowledge. We further propose an attention-based reasoning approach to model the intra-personal emotion contexts, inter-personal emotion propagation, and the personalities of different individuals. Experimental results demonstrate that our proposed approach outperforms related baselines significantly. We isolate and analyze the validity of different reasoning modules across various emotions of speakers and non-speakers. Finally, we draw forth several future research directions for multimodal emotion reasoning with MEmoR, aiming to empower high Emotional Quotient (EQ) in modern artificial intelligence systems. The code and dataset released on https://github.com/sunlightsgy/MEmoR.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- AVQA: A Dataset for Audio-Visual Question Answering on VideosPinci Yang, Xin Wang, Xuguang Duan, Hong Chen 等ACM MM 2022 · 被引用 60 次
- Aesthetic-Aware Image Style TransferZhiyuan Hu, Jia Jia, Bei Liu, Yaohua Bu 等ACM MM 2020 · 被引用 39 次
- IMOL: Incomplete-Modality-Tolerant Learning for Multi-Domain Fake News Video DetectionZhi Zeng, Jiaying Wu, Minnan Luo, Herun Wan 等ACL 2025 · 被引用 17 次
- Beyond Emotion Recognition: A Multi-Turn Multimodal Emotion Understanding and Reasoning BenchmarkJinpeng Hu, Hongchang Shi, Chongyuan Dai, Zhuo Li 等ACM MM 2025 · 被引用 7 次
- Hardness-Aware Dynamic Curriculum Learning for Robust Multimodal Emotion Recognition with Missing ModalitiesRui Liu, Haolin Zuo, Zheng Lian, Hongyu Yuan 等ACM MM 2025 · 被引用 6 次
它引用的顶会 Paper3
- CLEVRER: Collision Events for Video Representation and ReasoningKexin Yi, Chuang Gan, Yunzhu Li, Pushmeet Kohli 等ICLR 2020 · 被引用 584 次
- Boosting Visual Question Answering with Context-aware Knowledge AggregationGuohao Li, Xin Wang, Wenwu ZhuACM MM 2020 · 被引用 82 次
- Aesthetic-Aware Image Style TransferZhiyuan Hu, Jia Jia, Bei Liu, Yaohua Bu 等ACM MM 2020 · 被引用 39 次
相关 Paper
- iMiGUE: An Identity-Free Video Dataset for Micro-Gesture Understanding and Emotion AnalysisXin Liu, Henglin Shi, Haoyu Chen, Zitong Yu 等CVPR 2021
- DEEMO: De-identity Multimodal Emotion Recognition and ReasoningDeng Li, Bohao Xing, Xin Liu, Baiqiang Xia 等ACM MM 2025 · 被引用 8 次
- Pairwise Emotional Relationship Recognition in Drama Videos: Dataset and BenchmarkXun Gao, Yin Zhao, Jie Zhang, Longjun CaiACM MM 2021 · 被引用 8 次
- From Subtle Hints to Grand Expressions - Mastering Fine-grained Emotions with Dynamic Multimodal AnalysisQinfu Xu, Liyuan Pan, Shaozu Yuan, Yiwei Wei 等ACM MM 2025
- MIntRec: A New Dataset for Multimodal Intent RecognitionHanlei Zhang, Hua Xu, Xin Wang, Qianrui Zhou 等ACM MM 2022 · 被引用 66 次
