Open-Book Video Captioning With Retrieve-Copy-Generate Network
Ziqi Zhang, Zhongang Qi, Chunfeng Yuan, Ying Shan, Bing Li, Ying Deng, Weiming Hu
Abstract
In this paper, we convert traditional video captioning task into a new paradigm, i.e., Open-book Video Captioning, which generates natural language under the prompts of video-content-relevant sentences, not limited to the video itself. To address the open-book video captioning problem, we propose a novel Retrieve-Copy-Generate network, where a pluggable video-to-text retriever is constructed to retrieve sentences as hints from the training corpus effectively, and a copy-mechanism generator is introduced to extract expressions from multi-retrieved sentences dynamically. The two modules can be trained end-to-end or separately, which is flexible and extensible. Our framework coordinates the conventional retrieval-based methods with orthodox encoder-decoder methods, which can not only draw on the diverse expressions in the retrieved sentences but also generate natural and accurate content of the video. Extensive experiments on several benchmark datasets show that our proposed approach surpasses the state-of-the-art performance, indicating the effectiveness and promising of the proposed paradigm in the task of video captioning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3452f4f8-5684-470d-9fbd-3c6b82930aadCited by top-tier papers15
- SwinBERT: End-to-End Transformers with Sparse Attention for Video CaptioningKevin Lin, Linjie Li, Chung-Ching Lin, Faisal Ahmed et al.CVPR 2022 · 263 citations
- Expectation-Maximization Contrastive Learning for Compact Video-and-Language RepresentationsPeng Jin, Jinfa Huang, Fenglin Liu, Xian Wu et al.NeurIPS 2022 · 105 citations
- Accurate and Fast Compressed Video CaptioningYaojie Shen, Xin Gu, Kai Xu, Heng Fan et al.ICCV 2023 · 53 citations
- Learnability Matters: Active Learning for Video CaptioningYiqian Zhang, Buyu Liu, Jun Bao, Qiang Huang et al.NeurIPS 2024 · 47 citations
- OmniViD: A Generative Framework for Universal Video UnderstandingJunke Wang, Dongdong Chen, Chong Luo, Bo He et al.CVPR 2024 · 18 citations
Builds on14
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video ClipsAntoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi et al.ICCV 2019 · 1,437 citations
- VaTeX: A Large-Scale, High-Quality Multilingual Dataset for Video-and-Language ResearchXin Wang, Jiawei Wu, Jun-Kun Chen, Lei Li et al.ICCV 2019 · 688 citations
- Poly-encoders: Architectures and Pre-training Strategies for Fast and Accurate Multi-sentence ScoringSamuel Humeau, Kurt Shuster, Marie-Anne Lachaux, Jason WestonICLR 2020 · 316 citations
- Controllable Video Captioning With POS Sequence Guidance Based on Gated Fusion NetworkBairui Wang, Lin Ma, Wei Zhang, Wenhao Jiang et al.ICCV 2019 · 183 citations
Related papers
- Hierarchical Modular Network for Video CaptioningHanhua Ye, Guorong Li, Yuankai Qi, Shuhui Wang et al.CVPR 2022 · 95 citations
- Controllable Video Captioning with an Exemplar SentenceYitian Yuan, Lin Ma, Jingwen Wang, Wenwu ZhuACM MM 2020 · 21 citations
- Support-set bottlenecks for video-text representation learningMandela Patrick, Po-Yao Huang, Yuki Markus Asano, Florian Metze et al.ICLR 2021 · 269 citations
- Joint Syntax Representation Learning and Visual Cue Translation for Video CaptioningJingyi Hou, Xinxiao Wu, Wentian Zhao, Jiebo Luo et al.ICCV 2019 · 84 citations
- Leveraging Weighted Cross-Graph Attention for Visual and Semantic Enhanced Video Captioning NetworkDeepali Verma, Arya Haldar, Tanima DuttaAAAI 2023 · 13 citations
