MEmoR: A Dataset for Multimodal Emotion Reasoning in Videos
Guangyao Shen, Xin Wang, Xuguang Duan, Hongzhi Li, Wenwu Zhu
Abstract
Humans can perceive subtle emotions from various cues and contexts, even without hearing or seeing others. However, existing video datasets mainly focus on recognizing the emotions of the speakers from complete modalities. In this work, we present the task of multimodal emotion reasoning in videos. Beyond directly recognizing emotions from multimodal signals of target persons, this task requires a machine capable of reasoning about human emotions from the contexts and surrounding world. To facilitate the study towards this task, we introduce a new dataset, MEmoR, that provides fine-grained emotion annotations for both speakers and non-speakers. The videos in MEmoR are collected from TV shows closely in real-life scenarios. In these videos, while speakers may be non-visually described, non-speakers always deliver no audio-textual signals and are often visually inconspicuous. This modality-missing characteristic makes MEmoR a more practical yet challenging testbed for multimodal emotion reasoning. In support of various reasoning behaviors, the proposed MEmoR dataset provides both short-term contexts and external knowledge. We further propose an attention-based reasoning approach to model the intra-personal emotion contexts, inter-personal emotion propagation, and the personalities of different individuals. Experimental results demonstrate that our proposed approach outperforms related baselines significantly. We isolate and analyze the validity of different reasoning modules across various emotions of speakers and non-speakers. Finally, we draw forth several future research directions for multimodal emotion reasoning with MEmoR, aiming to empower high Emotional Quotient (EQ) in modern artificial intelligence systems. The code and dataset released on https://github.com/sunlightsgy/MEmoR.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6e840cfa-0a97-47b4-a069-5bd8167b1553Cited by top-tier papers7
- AVQA: A Dataset for Audio-Visual Question Answering on VideosPinci Yang, Xin Wang, Xuguang Duan, Hong Chen et al.ACM MM 2022 · 60 citations
- Aesthetic-Aware Image Style TransferZhiyuan Hu, Jia Jia, Bei Liu, Yaohua Bu et al.ACM MM 2020 · 39 citations
- IMOL: Incomplete-Modality-Tolerant Learning for Multi-Domain Fake News Video DetectionZhi Zeng, Jiaying Wu, Minnan Luo, Herun Wan et al.ACL 2025 · 17 citations
- Beyond Emotion Recognition: A Multi-Turn Multimodal Emotion Understanding and Reasoning BenchmarkJinpeng Hu, Hongchang Shi, Chongyuan Dai, Zhuo Li et al.ACM MM 2025 · 7 citations
- Hardness-Aware Dynamic Curriculum Learning for Robust Multimodal Emotion Recognition with Missing ModalitiesRui Liu, Haolin Zuo, Zheng Lian, Hongyu Yuan et al.ACM MM 2025 · 6 citations
Builds on3
- CLEVRER: Collision Events for Video Representation and ReasoningKexin Yi, Chuang Gan, Yunzhu Li, Pushmeet Kohli et al.ICLR 2020 · 584 citations
- Boosting Visual Question Answering with Context-aware Knowledge AggregationGuohao Li, Xin Wang, Wenwu ZhuACM MM 2020 · 82 citations
- Aesthetic-Aware Image Style TransferZhiyuan Hu, Jia Jia, Bei Liu, Yaohua Bu et al.ACM MM 2020 · 39 citations
Related papers
- iMiGUE: An Identity-Free Video Dataset for Micro-Gesture Understanding and Emotion AnalysisXin Liu, Henglin Shi, Haoyu Chen, Zitong Yu et al.CVPR 2021
- DEEMO: De-identity Multimodal Emotion Recognition and ReasoningDeng Li, Bohao Xing, Xin Liu, Baiqiang Xia et al.ACM MM 2025 · 8 citations
- Pairwise Emotional Relationship Recognition in Drama Videos: Dataset and BenchmarkXun Gao, Yin Zhao, Jie Zhang, Longjun CaiACM MM 2021 · 8 citations
- From Subtle Hints to Grand Expressions - Mastering Fine-grained Emotions with Dynamic Multimodal AnalysisQinfu Xu, Liyuan Pan, Shaozu Yuan, Yiwei Wei et al.ACM MM 2025
- MIntRec: A New Dataset for Multimodal Intent RecognitionHanlei Zhang, Hua Xu, Xin Wang, Qianrui Zhou et al.ACM MM 2022 · 66 citations
