Controllable Video Captioning with an Exemplar Sentence
Yitian Yuan, Lin Ma, Jingwen Wang, Wenwu Zhu
Abstract
In this paper, we investigate a novel and challenging task, namely controllable video captioning with an exemplar sentence. Formally, given a video and a syntactically valid exemplar sentence, the task aims to generate one caption which not only describes the semantic contents of the video, but also follows the syntactic form of the given exemplar sentence. In order to tackle such an exemplar-based video captioning task, we propose a novel Syntax Modulated Caption Generator (SMCG) incorporated in an encoder-decoder-reconstructor architecture. The proposed SMCG takes video semantic representation as an input, and conditionally modulates the gates and cells of long short-term memory network with respect to the encoded syntactic information of the given exemplar sentence. Therefore, SMCG is able to control the states for word prediction and achieve the syntax customized caption generation. We conduct experiments by collecting auxiliary exemplar sentences for two public video captioning datasets. Extensive experimental results demonstrate the effectiveness of our approach on generating syntax controllable and semantic preserved video captions. By providing different exemplar sentences, our approach is capable of producing different captions with various syntactic structures, thus indicating a promising way to strengthen the diversity of video captioning. Code for this paper is available at https://github.com/yytzsy/SMCG.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3ed9af61-c129-48ce-ba58-ae1f5569ab7eCited by top-tier papers6
- Latent Memory-augmented Graph Transformer for Visual StorytellingMengshi Qi, Jie Qin, Di Huang, Zhiqiang Shen et al.ACM MM 2021 · 18 citations
- Long-term Leap Attention, Short-term Periodic Shift for Video ClassificationHao Zhang, Lechao Cheng, Yanbin Hao, Chong-Wah NgoACM MM 2022 · 12 citations
- Set Prediction Guided by Semantic Concepts for Diverse Video CaptioningYifan Lu, Ziqi Zhang, Chunfeng Yuan, Peng Li et al.AAAI 2024 · 7 citations
- Edit As You Wish: Video Caption Editing with Multi-grained User ControlLinli Yao, Yuanmeng Zhang, Ziheng Wang, Xinglin Hou et al.ACM MM 2024 · 4 citations
- Open-Book Video Captioning With Retrieve-Copy-Generate NetworkZiqi Zhang, Zhongang Qi, Chunfeng Yuan, Ying Shan et al.CVPR 2021
Builds on1
Related papers
- Discriminative Latent Semantic Graph for Video CaptioningYang Bai, Junyan Wang, Yang Long, Bingzhang Hu et al.ACM MM 2021 · 26 citations
- Multi-modal Dependency Tree for Video CaptioningWentian Zhao, Xinxiao Wu, Jiebo LuoNeurIPS 2021 · 27 citations
- Joint Syntax Representation Learning and Visual Cue Translation for Video CaptioningJingyi Hou, Xinxiao Wu, Wentian Zhao, Jiebo Luo et al.ICCV 2019 · 84 citations
- Hierarchical Modular Network for Video CaptioningHanhua Ye, Guorong Li, Yuankai Qi, Shuhui Wang et al.CVPR 2022 · 95 citations
- Text with Knowledge Graph Augmented Transformer for Video CaptioningXin Gu, Guang Chen, Yufei Wang, Libo Zhang et al.CVPR 2023
