EmoVid: A Multimodal Emotion Video Dataset for Emotion-Centric Video Understanding and Generation
Zongyang Qiu, Bingyuan Wang, Xingbei Chen, Yingqing He, Zeyu Wang
摘要
Emotion plays a pivotal role in video-based expression, but existing video generation systems predominantly focus on low-level visual metrics while neglecting affective dimensions. Although emotion analysis has made progress in the visual domain, the video community lacks dedicated resources to bridge emotion understanding with generative tasks, particularly for stylized and non-realistic contexts. To address this gap, we introduce EmoVid, the first multimodal, emotion-annotated video dataset specifically designed for artistic media, which includes cartoon animations, movie clips, and animated stickers. Each video is annotated with emotion labels, visual attributes (brightness, colorfulness, hue), and text captions. Through systematic analysis, we uncover spatial and temporal patterns linking visual features to emotional perceptions across diverse video forms. Building on these insights, we develop an emotion-conditioned video generation technique by fine-tuning the Wan2.1 model. The results show a significant improvement in both quantitative metrics and the visual quality of generated videos for text-to-video and image-to-video tasks. EmoVid establishes a new benchmark and protocol for affective video computing. Our work not only offers valuable insights into visual emotion analysis in artistic videos but also provides practical methods for enhancing emotional expression in video generation. The extended version and the dataset are available on our project page.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper12
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific TuningYuwei Guo, Ceyuan Yang, Anyi Rao, Zhengyang Liang 等ICLR 2024 · 被引用 1,493 次
- CLIPScore: A Reference-free Evaluation Metric for Image CaptioningJack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras 等EMNLP 2021 · 被引用 937 次
- Context-Aware Emotion Recognition NetworksJiyoung Lee, Seungryong Kim, Sunok Kim, Jungin Park 等ICCV 2019 · 被引用 285 次
- EmoSet: A Large-scale Visual Emotion Dataset with Rich AttributesJingyuan Yang, Qirui Huang, Tingting Ding, Dani Lischinski 等ICCV 2023 · 被引用 111 次
相关 Paper
- EmoDETective: Detecting, Exploring, and Thinking Emotional Cause in VideosXuandong Huang, Yuzhe Zhou, Jiashu Li, Shiqian Lu 等ACM MM 2025 · 被引用 1 次
- Observe before Generate: Emotion-Cause aware Video Caption for Multimodal Emotion Cause Generation in ConversationsFanfan Wang, Heqing Ma, Xiangqing Shen, Jianfei Yu 等ACM MM 2024 · 被引用 6 次
- MEmoR: A Dataset for Multimodal Emotion Reasoning in VideosGuangyao Shen, Xin Wang, Xuguang Duan, Hongzhi Li 等ACM MM 2020 · 被引用 38 次
- Make Me Happier: Evoking Emotions through Image Diffusion ModelsQing Lin, Jingfeng Zhang, Yew-Soon Ong, Mengmi ZhangICCV 2025 · 被引用 4 次
- V2C: Visual Voice CloningQi Chen, Mingkui Tan, Yuankai Qi, Jiaqiu Zhou 等CVPR 2022
