Syntax-Aware Action Targeting for Video Captioning
Qi Zheng, Chaoyue Wang, Dacheng Tao
摘要
Video captioning aims to describe objects and their interactions in the video using natural language. Existing methods have made great efforts to identify objects in videos, but few of them emphasize the prediction of interactions among objects, which is usually indicated by action/predicate in generated sentences. Different from other components in a sentence, the predicate depends on both the static scene and the dynamic motions in a video. Due to the neglect of such uniqueness, actions generated by existing methods may depend heavily on the co-occurrence of objects, e.g. 'driving' is predicted with high confidence whenever both man and car are detected. In this paper, we propose a Syntax-Aware Action Targeting (SAAT) module that explicitly learns actions by simultaneously referring to the subject and video dynamics. Specifically, we first identify the subject by drawing global dependence among multiple objects, and then decode action from a common space that fuses the embedding of the subject and the temporal feature of the video. Validated on two public datasets, the proposed module increases action accuracy in generated descriptions, which present better semantic consistency with the dynamic content in videos. Codes are available on https://github.com/SydCaption/SAAT .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- SwinBERT: End-to-End Transformers with Sparse Attention for Video CaptioningKevin Lin, Linjie Li, Chung-Ching Lin, Faisal Ahmed 等CVPR 2022 · 被引用 263 次
- Semantic Grouping Network for Video CaptioningHobin Ryu, Sunghun Kang, Haeyong Kang, Chang D. YooAAAI 2021 · 被引用 160 次
- Hierarchical Modular Network for Video CaptioningHanhua Ye, Guorong Li, Yuankai Qi, Shuhui Wang 等CVPR 2022 · 被引用 95 次
- Motion Guided Region Message Passing for Video CaptioningShaoxiang Chen, Yu-Gang JiangICCV 2021 · 被引用 71 次
- Refined Semantic Enhancement towards Frequency Diffusion for Video CaptioningXian Zhong, Zipeng Li, Shuqin Chen, Kui Jiang 等AAAI 2023 · 被引用 70 次
它引用的顶会 Paper1
相关 Paper
- Interventional Video Relation DetectionYicong Li, Xun Yang, Xindi Shang, Tat-Seng ChuaACM MM 2021 · 被引用 61 次
- Comprehensive Visual Grounding for Video DescriptionWenhui Jiang, Yibo Cheng, Linxin Liu, Yuming Fang 等AAAI 2024 · 被引用 5 次
- Discriminative Latent Semantic Graph for Video CaptioningYang Bai, Junyan Wang, Yang Long, Bingzhang Hu 等ACM MM 2021 · 被引用 26 次
- Joint Syntax Representation Learning and Visual Cue Translation for Video CaptioningJingyi Hou, Xinxiao Wu, Wentian Zhao, Jiebo Luo 等ICCV 2019 · 被引用 84 次
- Set Prediction Guided by Semantic Concepts for Diverse Video CaptioningYifan Lu, Ziqi Zhang, Chunfeng Yuan, Peng Li 等AAAI 2024 · 被引用 7 次
