Observe before Generate: Emotion-Cause aware Video Caption for Multimodal Emotion Cause Generation in Conversations
Fanfan Wang, Heqing Ma, Xiangqing Shen, Jianfei Yu, Rui Xia
Abstract
Emotion cause analysis has attracted increasing attention in recent years. However, the integration of multimodal information with emotion causes remains underexplored. Existing studies merely extract utterances from conversations as cause evidence, which is too coarse-grained to locate the exact causes from other modalities, especially those that may be reflected only in a specific video frame of an utterance. To address these limitations, we introduce a new task named Multimodal Emotion Cause Generation in Conversations (MECGC), which aims to generate an abstractive summary clearly and intuitively describing the causes that trigger the given emotion based on the multimodal context of conversations. We accordingly construct a dataset named ECGF that contains 1,374 conversations and 7,690 emotion instances from TV series. We further develop a generative framework that first generates emotion-cause aware video captions (Observe) and then facilitates the generation of emotion causes (Generate). The captioning model is trained with examples synthesized by a Multimodal Large Language Model (MLLM). Experimental results demonstrate the effectiveness of our framework and the significance of multimodal information for emotion cause analysis. Our dataset and source codes are available at https://github.com/NUSTM/MECGC.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get cfc0a9f0-8520-47c0-94c5-bdf255e69a2dCited by top-tier papers3
- Emotion across Modalities and Cultures: Multilingual Multimodal Emotion-Cause Analysis with Memory-inspired FrameworkDan Wu, Xincheng Ju, Dong Zhang, Shoushan Li et al.ACM MM 2025 · 2 citations
- Voices, Faces, and Feelings: Multi-modal Emotion-Cognition Captioning for Mental Health UnderstandingZhiyuan Zhou, Yanrong Guo, Shijie HaoAAAI 2026
- Locate and Explain: Joint Multimodal Emotion Cause Extraction and Summarization in ConversationJikun Wan, Chen Gong, Guohong FuACL 2026
Related papers
- EmoDETective: Detecting, Exploring, and Thinking Emotional Cause in VideosXuandong Huang, Yuzhe Zhou, Jiashu Li, Shiqian Lu et al.ACM MM 2025 · 1 citation
- AffectGPT: A New Dataset, Model, and Benchmark for Emotion Understanding with Multimodal Large Language ModelsZheng Lian, Haoyu Chen, Lan Chen, Haiyang Sun et al.ICML 2025
- ECFCON: Emotion Consequence Forecasting in ConversationsXincheng Ju, Dong Zhang, Suyang Zhu, Junhui Li et al.ACM MM 2024 · 1 citation
- Causal-ERC: A Multimodal Framework with Causal Prompting for Emotion Recognition in Conversations with Large Language ModelsRan Jing, Geng Tu, Yice Zhang, Ruifeng XuAAAI 2026
- ECERC: Evidence-Cause Attention Network for Multi-Modal Emotion Recognition in ConversationTao Zhang, Zhenhua TanACL 2025 · 4 citations
