Semantic Grouping Network for Video Captioning
Hobin Ryu, Sunghun Kang, Haeyong Kang, Chang D. Yoo
摘要
This paper considers a video caption generating network referred to as Semantic Grouping Network (SGN) that attempts (1) to group video frames with discriminating word phrases of partially decoded caption and then (2) to decode those semantically aligned groups in predicting the next word. As consecutive frames are not likely to provide unique information, prior methods have focused on discarding or merging repetitive information based only on the input video. The SGN learns an algorithm to capture the most discriminating word phrases of the partially decoded caption and a mapping that associates each phrase to the relevant video frames - establishing this mapping allows semantically related frames to be clustered, which reduces redundancy. In contrast to the prior methods, the continuous feedback from decoded words enables the SGN to dynamically update the video representation that adapts to the partially decoded caption. Furthermore, a contrastive attention loss is proposed to facilitate accurate alignment between a word phrase and video frames without manual annotations. The SGN achieves state-of-the-art performances by outperforming runner-up methods by a margin of 2.1%p and 2.4%p in a CIDEr-D score on MSVD and MSR-VTT datasets, respectively. Extensive experiments demonstrate the effectiveness and interpretability of the SGN.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- Hierarchical Modular Network for Video CaptioningHanhua Ye, Guorong Li, Yuankai Qi, Shuhui Wang 等CVPR 2022 · 被引用 95 次
- Refined Semantic Enhancement towards Frequency Diffusion for Video CaptioningXian Zhong, Zipeng Li, Shuqin Chen, Kui Jiang 等AAAI 2023 · 被引用 70 次
- Accurate and Fast Compressed Video CaptioningYaojie Shen, Xin Gu, Kai Xu, Heng Fan 等ICCV 2023 · 被引用 53 次
- Phrase-Level Temporal Relationship Mining for Temporal Sentence LocalizationMinghang Zheng, Sizhe Li, Qingchao Chen, Yuxin Peng 等AAAI 2023 · 被引用 26 次
- Towards Surveillance Video-and-Language Understanding: New Dataset, Baselines, and ChallengesTongtong Yuan, Xuange Zhang, Kun Liu, Bo Liu 等CVPR 2024 · 被引用 23 次
它引用的顶会 Paper5
- Memory Enhanced Global-Local Aggregation for Video Object DetectionYihong Chen, Yue Cao, Han Hu, Liwei WangCVPR 2020
- Spatio-Temporal Graph for Video Captioning With Knowledge DistillationBoxiao Pan, Haoye Cai, De-An Huang, Kuan-Hui Lee 等CVPR 2020
- Object Relational Graph With Teacher-Recommended Learning for Video CaptioningZiqi Zhang, Yaya Shi, Chunfeng Yuan, Bing Li 等CVPR 2020
- Modality Shifting Attention Network for Multi-Modal Video Question AnsweringJunyeong Kim, Minuk Ma, Trung X. Pham, Kyungsu Kim 等CVPR 2020
- Syntax-Aware Action Targeting for Video CaptioningQi Zheng, Chaoyue Wang, Dacheng TaoCVPR 2020
相关 Paper
- Discriminative Latent Semantic Graph for Video CaptioningYang Bai, Junyan Wang, Yang Long, Bingzhang Hu 等ACM MM 2021 · 被引用 26 次
- Leveraging Weighted Cross-Graph Attention for Visual and Semantic Enhanced Video Captioning NetworkDeepali Verma, Arya Haldar, Tanima DuttaAAAI 2023 · 被引用 13 次
- Non-Autoregressive Coarse-to-Fine Video CaptioningBang Yang, Yuexian Zou, Fenglin Liu, Can ZhangAAAI 2021 · 被引用 92 次
- Set Prediction Guided by Semantic Concepts for Diverse Video CaptioningYifan Lu, Ziqi Zhang, Chunfeng Yuan, Peng Li 等AAAI 2024 · 被引用 7 次
- Image Captioning with Context-Aware Auxiliary GuidanceZeliang Song, Xiaofei Zhou, Zhendong Mao, Jianlong TanAAAI 2021 · 被引用 36 次
