Emotion-Prior Awareness Network for Emotional Video Captioning
Peipei Song, Dan Guo, Xun Yang, Shengeng Tang, Erkun Yang, Meng Wang
摘要
Emotional video captioning (EVC) is an emerging task to describe the factual content with the inherent emotion expressed in a video. It is crucial for the EVC task to effectively perceive subtle and ambiguous visual emotion cues in the stage of caption generation. However, existing captioning methods usually overlooked the learning of emotions in user-generated videos, thus making the generated sentence a bit boring and soulless.
To address this issue, this paper proposes a new emotional captioning perspective in a human-like perception-priority manner, i.e., first perceiving the inherent emotion and then leveraging the perceived emotion cue to support caption generation. Specifically, we devise an Emotion-Prior Awareness Network (EPAN). It mainly benefits from a novel tree-structured emotion learning module involving both catalog-level psychological categories and lexical-level usual words to achieve the goal of explicit and fine-grained emotion perception. Besides, we develop a novel subordinate emotion masking mechanism between the catalog level and lexical level that facilitates coarse-to-fine emotion learning. Afterward, with the emotion prior, we can effectively decode the emotional caption by exploiting the complementation of visual, textual, and emotional semantics. In addition, we also introduce three simple yet effective optimization objectives, which can significantly boost the emotion learning from the perspectives of emotional captioning, hierarchical emotion classification, and emotional contrastive learning. Sufficient experimental results on three benchmark datasets clearly demonstrate the advantages of our proposed EPAN over existing
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Enhancing One-Shot Federated Learning Through Data and Ensemble Co-BoostingRong Dai, Yonggang Zhang, Ang Li, Tongliang Liu 等ICLR 2024 · 被引用 40 次
- Sign-IDD: Iconicity Disentangled Diffusion for Sign Language ProductionShengeng Tang, Jiayi He, Dan Guo, Yanyan Wei 等AAAI 2025 · 被引用 23 次
- Dual-path Collaborative Generation Network for Emotional Video CaptioningCheng Ye, Weidong Chen, Jingyu Li, Lei Zhang 等ACM MM 2024 · 被引用 15 次
- Towards Understanding Future: Consistency Guided Probabilistic Modeling for Action AnticipationZhao Xie, Yadong Shi, Kewei Wu, Yaru Cheng 等AAAI 2024 · 被引用 9 次
- Benchmarking and Bridging Emotion Conflicts for Multimodal Emotion ReasoningZhiyuan Han, Beier Zhu, Yanlong Xu, Peipei Song 等ACM MM 2025 · 被引用 7 次
它引用的顶会 Paper22
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Deconfounded Video Moment Retrieval with Causal InterventionXun Yang, Fuli Feng, Wei Ji, Meng Wang 等SIGIR 2021 · 被引用 198 次
- Proposal-Free Video Grounding with Contextual Pyramid NetworkKun Li, Dan Guo, Meng WangAAAI 2021 · 被引用 138 次
- Tree-Augmented Cross-Modal Encoding for Complex-Query Video RetrievalXun Yang, Jianfeng Dong, Yixin Cao, Xun Wang 等SIGIR 2020 · 被引用 131 次
- Knowledge Bridging for Empathetic Dialogue GenerationQintong Li, Piji Li, Zhaochun Ren, Pengjie Ren 等AAAI 2022 · 被引用 128 次
相关 Paper
- Multi-round Mutual Emotion-Cause Pair Extraction for Emotion-Attributed Video CaptioningCheng Ye, Weidong Chen, Peipei Song, Xinyan Liu 等ACM MM 2025 · 被引用 9 次
- Progressive Visual Content Understanding Network for Image Emotion ClassificationJicai Pan, Shangfei WangACM MM 2023 · 被引用 5 次
- An End-to-End Visual-Audio Attention Network for Emotion Recognition in User-Generated VideosSicheng Zhao, Yunsheng Ma, Yang Gu, Jufeng Yang 等AAAI 2020 · 被引用 123 次
- Hierarchical Context-aware Network for Dense Video Event CaptioningLei Ji, Xianglin Guo, Haoyang Huang, Xilin ChenACL 2021
- Multi-Perspective Video CaptioningYi Bin, Xindi Shang, Bo Peng, Yujuan Ding 等ACM MM 2021 · 被引用 14 次
