Relational Graph Learning for Grounded Video Description Generation
Wenqiao Zhang, Xin Eric Wang, Siliang Tang, Haizhou Shi, Haochen Shi, Jun Xiao, Yueting Zhuang, William Yang Wang
摘要
Grounded video description (GVD) encourages captioning models to attend to appropriate video regions (e.g., objects) dynamically and generate a description. Such a setting can help explain the decisions of captioning models and prevents the model from hallucinating object words in its description. However, such design mainly focuses on object word generation and thus may ignore fine-grained information and suffer from missing visual concepts. Moreover, relational words (e.g., 'jump left or right') are usual spatio-temporal inference results, i.e., these words cannot be grounded on certain spatial regions. To tackle the above limitations, we design a novel relational graph learning framework for GVD, in which a language-refined scene graph representation is designed to explore fine-grained visual concepts. Furthermore, the refined graph can be regarded as relational inductive knowledge to assist captioning models in selecting the relevant information it needs to generate correct words. We validate the effectiveness of our model through automatic metrics and human evaluation, and the results indicate that our approach can generate more fine-grained and accurate description, and it solves the problem of object hallucination to some extent.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- BoostMIS: Boosting Medical Image Semi-supervised Learning with Adaptive Pseudo Labeling and Informative Active AnnotationWenqiao Zhang, Lei Zhu, James Hallinan, Shengyu Zhang 等CVPR 2022 · 被引用 115 次
- Re4: Learning to Re-contrast, Re-attend, Re-construct for Multi-interest RecommendationShengyu Zhang, Lingxiao Yang, Dong Yao, Yujie Lu 等WWW 2022 · 被引用 67 次
- Consensus Graph Representation Learning for Better Grounded Image CaptioningWenqiao Zhang, Haochen Shi, Siliang Tang, Jun Xiao 等AAAI 2021 · 被引用 63 次
- Compositional Temporal Grounding with Structured Variational Cross-Graph Correspondence LearningJuncheng Li, Junlin Xie, Long Qian, Linchao Zhu 等CVPR 2022 · 被引用 63 次
- Semantic-Aware Representation Blending for Multi-Label Image Recognition with Partial LabelsTao Pu, Tianshui Chen, Hefeng Wu, Liang LinAAAI 2022 · 被引用 58 次
它引用的顶会 Paper4
- Boundary-Aware Feature Propagation for Scene SegmentationHenghui Ding, Xudong Jiang, Ai Qun Liu, Nadia Magnenat-Thalmann 等ICCV 2019 · 被引用 283 次
- Unpaired Image Captioning via Scene Graph AlignmentsJiuxiang Gu, Shafiq R. Joty, Jianfei Cai, Handong Zhao 等ICCV 2019 · 被引用 191 次
- Consensus Graph Representation Learning for Better Grounded Image CaptioningWenqiao Zhang, Haochen Shi, Siliang Tang, Jun Xiao 等AAAI 2021 · 被引用 63 次
- Photo Stream Question AnswerWenqiao Zhang, Siliang Tang, Yanpeng Cao, Jun Xiao 等ACM MM 2020 · 被引用 5 次
相关 Paper
- Object Relational Graph With Teacher-Recommended Learning for Video CaptioningZiqi Zhang, Yaya Shi, Chunfeng Yuan, Bing Li 等CVPR 2020
- Multimodal Graph Conditioned Diffusion Model for Video CaptioningBenhui Zhang, Junyu Gao, Yuan YuanWWW 2026
- Weakly-Supervised Generation and Grounding of Visual Descriptions with Conditional Generative ModelsEffrosyni Mavroudi, René VidalCVPR 2022 · 被引用 5 次
- Discriminative Latent Semantic Graph for Video CaptioningYang Bai, Junyan Wang, Yang Long, Bingzhang Hu 等ACM MM 2021 · 被引用 26 次
- Large-Scale Pre-Training for Grounded Video Caption GenerationEvangelos Kazakos, Cordelia Schmid, Josef SivicICCV 2025
