How You Feelin'? Learning Emotions and Mental States in Movie Scenes
Dhruv Srivastava, Aditya Kumar Singh, Makarand Tapaswi
摘要
Movie story analysis requires understanding characters' emotions and mental states. Towards this goal, we formulate emotion understanding as predicting a diverse and multi-label set of emotions at the level of a movie scene and for each character. We propose EmoTx, a multimodal Transformer-based architecture that ingests videos, multiple characters, and dialog utterances to make joint predictions. By leveraging annotations from the MovieGraphs dataset [72], we aim to predict classic emotions (e.g. happy, angry) and other mental states (e.g. honest, helpful). We conduct experiments on the most frequently occurring 10 and 25 labels, and a mapping that clusters 181 labels to 26. Ablation studies and comparison against adapted stateof-the-art emotion recognition approaches shows the effectiveness of EmoTx. Analyzing EmoTx's self-attention scores reveals that expressive emotions often look at character tokens while other mental states rely on video and dialog cues.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Two in One Go: Single-stage Emotion Recognition with Decoupled Subject-context TransformerXinpeng Li, Teng Wang, Jian Zhao, Shuyi Mao 等ACM MM 2024 · 被引用 3 次
- "Previously on..." from Recaps to Story SummarizationAditya Kumar Singh, Dhruv Srivastava, Makarand TapaswiCVPR 2024 · 被引用 1 次
它引用的顶会 Paper17
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Multiscale Vision TransformersHaoqi Fan, Bo Xiong, Karttikeya Mangalam, Yanghao Li 等ICCV 2021 · 被引用 1,611 次
- VideoBERT: A Joint Model for Video and Language Representation LearningChen Sun, Austin Myers, Carl Vondrick, Kevin Murphy 等ICCV 2019 · 被引用 1,396 次
- Context-Aware Emotion Recognition NetworksJiyoung Lee, Seungryong Kim, Sunok Kim, Jungin Park 等ICCV 2019 · 被引用 285 次
- DialogXL: All-in-One XLNet for Multi-Party Conversation Emotion RecognitionWeizhou Shen, Junqing Chen, Xiaojun Quan, Zhixian XieAAAI 2021 · 被引用 251 次
相关 Paper
- Modeling Human Motives and Emotions from Personal Narratives Using External Knowledge And Entity TrackingPrashanth Vijayaraghavan, Deb RoyWWW 2021 · 被引用 10 次
- Transformer-based Label Set Generation for Multi-modal Multi-label Emotion DetectionXincheng Ju, Dong Zhang, Junhui Li, Guodong ZhouACM MM 2020 · 被引用 64 次
- Affect2MM: Affective Analysis of Multimedia Content Using Emotion CausalityTrisha Mittal, Puneet Mathur, Aniket Bera, Dinesh ManochaCVPR 2021
- Observe before Generate: Emotion-Cause aware Video Caption for Multimodal Emotion Cause Generation in ConversationsFanfan Wang, Heqing Ma, Xiangqing Shen, Jianfei Yu 等ACM MM 2024 · 被引用 6 次
- M3ED: Multi-modal Multi-scene Multi-label Emotional Dialogue DatabaseJinming Zhao, Tenggan Zhang, Jingwen Hu, Yuchen Liu 等ACL 2022 · 被引用 88 次
