Multi-modal Dependency Tree for Video Captioning
Wentian Zhao, Xinxiao Wu, Jiebo Luo
Abstract
Generating fluent and relevant language to describe visual content is critical for the video captioning task. Many existing methods generate captions using sequence models that predict words in a left-to-right order. In this paper, we investigate a graph structured model by explicitly modeling the hierarchical structure in the sentences to further improve the fluency and relevance of the generated captions. To this end, we propose a novel video captioning method that generates a sentence by first constructing a multi-modal dependency tree and then traversing the constructed tree, where the syntactic structure and semantic relationship in the sentence are represented by the tree topology. To take full advantage of the information from both vision and language, both the visual and textual representation features are encoded into each tree node. Different from existing dependency parsing methods that generate uni-modal dependency trees for language understanding, our method constructs multi-modal dependency trees for language generation of videos. We also propose a tree-structured reinforcement learning algorithm to effectively optimize the captioning model, where a novel reward is designed by evaluating the semantic consistency between the generated sub-trees and the ground-truth tree. Extensive experiments on several video captioning datasets demonstrate the effectiveness of the proposed method.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Leveraging Weighted Cross-Graph Attention for Visual and Semantic Enhanced Video Captioning NetworkDeepali Verma, Arya Haldar, Tanima DuttaAAAI 2023 · 13 citations
- Learning the Dynamics of Visual Relational Reasoning via Reinforced Path RoutingChenchen Jing, Yunde Jia, Yuwei Wu, Chuanhao Li et al.AAAI 2022 · 5 citations
Builds on9
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- Attention on Attention for Image CaptioningLun Huang, Wenmin Wang, Jie Chen, Xiaoyong WeiICCV 2019 · 992 citations
- Entangled Transformer for Image CaptioningGuang Li, Linchao Zhu, Ping Liu, Yi YangICCV 2019 · 346 citations
- Controllable Video Captioning With POS Sequence Guidance Based on Gated Fusion NetworkBairui Wang, Lin Ma, Wei Zhang, Wenhao Jiang et al.ICCV 2019 · 183 citations
- Watch, Listen and Tell: Multi-Modal Weakly Supervised Dense Event CaptioningTanzila Rahman, Bicheng Xu, Leonid SigalICCV 2019 · 89 citations
Related papers
- Multimodal Graph Conditioned Diffusion Model for Video CaptioningBenhui Zhang, Junyu Gao, Yuan YuanWWW 2026
- Object Relational Graph With Teacher-Recommended Learning for Video CaptioningZiqi Zhang, Yaya Shi, Chunfeng Yuan, Bing Li et al.CVPR 2020
- Joint Syntax Representation Learning and Visual Cue Translation for Video CaptioningJingyi Hou, Xinxiao Wu, Wentian Zhao, Jiebo Luo et al.ICCV 2019 · 84 citations
- Triangle-Reward Reinforcement Learning: A Visual-Linguistic Semantic Alignment for Image CaptioningWeizhi Nie, Jiesi Li, Ning Xu, An-An Liu et al.ACM MM 2021 · 9 citations
- Controllable Video Captioning with an Exemplar SentenceYitian Yuan, Lin Ma, Jingwen Wang, Wenwu ZhuACM MM 2020 · 21 citations
