Relational Graph Learning for Grounded Video Description Generation
Wenqiao Zhang, Xin Eric Wang, Siliang Tang, Haizhou Shi, Haochen Shi, Jun Xiao, Yueting Zhuang, William Yang Wang
Abstract
Grounded video description (GVD) encourages captioning models to attend to appropriate video regions (e.g., objects) dynamically and generate a description. Such a setting can help explain the decisions of captioning models and prevents the model from hallucinating object words in its description. However, such design mainly focuses on object word generation and thus may ignore fine-grained information and suffer from missing visual concepts. Moreover, relational words (e.g., 'jump left or right') are usual spatio-temporal inference results, i.e., these words cannot be grounded on certain spatial regions. To tackle the above limitations, we design a novel relational graph learning framework for GVD, in which a language-refined scene graph representation is designed to explore fine-grained visual concepts. Furthermore, the refined graph can be regarded as relational inductive knowledge to assist captioning models in selecting the relevant information it needs to generate correct words. We validate the effectiveness of our model through automatic metrics and human evaluation, and the results indicate that our approach can generate more fine-grained and accurate description, and it solves the problem of object hallucination to some extent.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8227d6b7-62cc-4cf7-bf8b-294d4495fbc3Cited by top-tier papers14
- BoostMIS: Boosting Medical Image Semi-supervised Learning with Adaptive Pseudo Labeling and Informative Active AnnotationWenqiao Zhang, Lei Zhu, James Hallinan, Shengyu Zhang et al.CVPR 2022 · 115 citations
- Re4: Learning to Re-contrast, Re-attend, Re-construct for Multi-interest RecommendationShengyu Zhang, Lingxiao Yang, Dong Yao, Yujie Lu et al.WWW 2022 · 67 citations
- Consensus Graph Representation Learning for Better Grounded Image CaptioningWenqiao Zhang, Haochen Shi, Siliang Tang, Jun Xiao et al.AAAI 2021 · 63 citations
- Compositional Temporal Grounding with Structured Variational Cross-Graph Correspondence LearningJuncheng Li, Junlin Xie, Long Qian, Linchao Zhu et al.CVPR 2022 · 63 citations
- Semantic-Aware Representation Blending for Multi-Label Image Recognition with Partial LabelsTao Pu, Tianshui Chen, Hefeng Wu, Liang LinAAAI 2022 · 58 citations
Builds on4
- Boundary-Aware Feature Propagation for Scene SegmentationHenghui Ding, Xudong Jiang, Ai Qun Liu, Nadia Magnenat-Thalmann et al.ICCV 2019 · 283 citations
- Unpaired Image Captioning via Scene Graph AlignmentsJiuxiang Gu, Shafiq R. Joty, Jianfei Cai, Handong Zhao et al.ICCV 2019 · 191 citations
- Consensus Graph Representation Learning for Better Grounded Image CaptioningWenqiao Zhang, Haochen Shi, Siliang Tang, Jun Xiao et al.AAAI 2021 · 63 citations
- Photo Stream Question AnswerWenqiao Zhang, Siliang Tang, Yanpeng Cao, Jun Xiao et al.ACM MM 2020 · 5 citations
Related papers
- Object Relational Graph With Teacher-Recommended Learning for Video CaptioningZiqi Zhang, Yaya Shi, Chunfeng Yuan, Bing Li et al.CVPR 2020
- Multimodal Graph Conditioned Diffusion Model for Video CaptioningBenhui Zhang, Junyu Gao, Yuan YuanWWW 2026
- Weakly-Supervised Generation and Grounding of Visual Descriptions with Conditional Generative ModelsEffrosyni Mavroudi, René VidalCVPR 2022 · 5 citations
- Discriminative Latent Semantic Graph for Video CaptioningYang Bai, Junyan Wang, Yang Long, Bingzhang Hu et al.ACM MM 2021 · 26 citations
- Large-Scale Pre-Training for Grounded Video Caption GenerationEvangelos Kazakos, Cordelia Schmid, Josef SivicICCV 2025
