Dual-path Collaborative Generation Network for Emotional Video Captioning
Cheng Ye, Weidong Chen, Jingyu Li, Lei Zhang, Zhendong Mao
Abstract
Emotional Video Captioning (EVC) is an emerging task that aims to describe factual content with the intrinsic emotions expressed in videos. The essential of the EVC task is to effectively perceive subtle and ambiguous visual emotional cues during the caption generation, which is neglected by the traditional video captioning. Existing emotional video captioning methods perceive global visual emotional cues at first, and then combine them with the video features to guide the emotional caption generation, which neglects two characteristics of the EVC task. Firstly, their methods neglect the dynamic subtle changes in the intrinsic emotions of the video, which makes it difficult to meet the needs of common scenes with diverse and changeable emotions. Secondly, as their methods incorporate emotional cues into each step, the guidance role of emotion is overemphasized, which makes factual content more or less ignored during generation. To this end, we propose a dual-path collaborative generation network, which dynamically perceives visual emotional cues evolutions while generating emotional captions by collaborative learning. The two paths promote each other and significantly improve the generation performance. Specifically, in the dynamic emotion perception path, we propose a dynamic emotion evolution module, which first aggregates visual features and historical caption features to summarize the global visual emotional cues, and then dynamically selects emotional cues required to be re-composed at each stage as well as re-composed them to achieve emotion evolution by dynamically enhancing or suppressing different granularity subspace's semantics. Besides, in the adaptive caption generation path, to balance the description of factual content and emotional cues, we propose an emotion adaptive decoder, which firstly estimates emotion intensity via the alignment of emotional features and historical caption features at each generation step, and then, emotional guidance adaptively incorporate into the caption generation based on the emotional intensity. Thus, our methods can generate emotion-related words at the necessary time step, and our caption generation balances the guidance of factual content and emotional cues well. Extensive experiments on three challenging datasets demonstrate the superiority of our approach and each proposed module.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Voices, Faces, and Feelings: Multi-modal Emotion-Cognition Captioning for Mental Health UnderstandingZhiyuan Zhou, Yanrong Guo, Shijie HaoAAAI 2026
- Self-Critical Distillation Network for Video-based Commonsense CaptioningMengqi Yuan, Gengyun Jia, Bing-Kun BaoCVPR 2026
Builds on23
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- LightGCN: Simplifying and Powering Graph Convolution Network for RecommendationXiangnan He, Kuan Deng, Xiang Wang, Yan Li et al.SIGIR 2020 · 4,448 citations
- VideoBERT: A Joint Model for Video and Language Representation LearningChen Sun, Austin Myers, Carl Vondrick, Kevin Murphy et al.ICCV 2019 · 1,396 citations
- Robust Lightweight Facial Expression Recognition Network with Label Distribution TrainingZengqun Zhao, Qingshan Liu, Feng ZhouAAAI 2021 · 300 citations
- An End-to-End Visual-Audio Attention Network for Emotion Recognition in User-Generated VideosSicheng Zhao, Yunsheng Ma, Yang Gu, Jufeng Yang et al.AAAI 2020 · 123 citations
Related papers
- Multi-round Mutual Emotion-Cause Pair Extraction for Emotion-Attributed Video CaptioningCheng Ye, Weidong Chen, Peipei Song, Xinyan Liu et al.ACM MM 2025 · 9 citations
- Emotion-Prior Awareness Network for Emotional Video CaptioningPeipei Song, Dan Guo, Xun Yang, Shengeng Tang et al.ACM MM 2023 · 29 citations
- Task-Specific Information Decomposition for End-to-End Dense Video CaptioningZhiyue Liu, Xinru Zhang, Jinyuan LiuACL 2025
- Hierarchical Context-aware Network for Dense Video Event CaptioningLei Ji, Xianglin Guo, Haoyang Huang, Xilin ChenACL 2021
- Implicit Location-Caption Alignment via Complementary Masking for Weakly-Supervised Dense Video CaptioningShiping Ge, Qiang Chen, Zhiwei Jiang, Yafeng Yin et al.AAAI 2025 · 7 citations
