Text with Knowledge Graph Augmented Transformer for Video Captioning
Xin Gu, Guang Chen, Yufei Wang, Libo Zhang, Tiejian Luo, Longyin Wen
摘要
Video captioning aims to describe the content of videos using natural language. Although significant progress has been made, there is still much room to improve the performance for real-world applications, mainly due to the long-tail words challenge. In this paper, we propose a text with knowledge graph augmented transformer (TextKG) for video captioning. Notably, TextKG is a two-stream transformer, formed by the external stream and internal stream. The external stream is designed to absorb additional knowledge, which models the interactions between the additional knowledge, e.g., pre-built knowledge graph, and the builtin information of videos, e.g., the salient object regions, speech transcripts, and video captions, to mitigate the longtail words challenge. Meanwhile, the internal stream is designed to exploit the multi-modality information in videos (e.g., the appearance of video frames, speech transcripts, and video captions) to ensure the quality of caption results. In addition, the cross attention mechanism is also used in between the two streams for sharing information. In this way, the two streams can help each other for more accurate results. Extensive experiments conducted on four challenging video captioning datasets, i.e., YouCookII, ActivityNet Captions, MSR-VTT, and MSVD, demonstrate that the proposed method performs favorably against the state-of-theart methods. Specifically, the proposed TextKG method outperforms the best published results by improving 18.7% absolute CIDEr scores on the YouCookII dataset.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- NExT-GPT: Any-to-Any Multimodal LLMShengqiong Wu, Hao Fei, Leigang Qu, Wei Ji 等ICML 2024 · 被引用 786 次
- Vitron: A Unified Pixel-level Vision LLM for Understanding, Generating, Segmenting, EditingHao Fei, Shengqiong Wu, Hanwang Zhang, Tat-Seng Chua 等NeurIPS 2024 · 被引用 100 次
- Accurate and Fast Compressed Video CaptioningYaojie Shen, Xin Gu, Kai Xu, Heng Fan 等ICCV 2023 · 被引用 53 次
- NativE: Multi-modal Knowledge Graph Completion in the WildYichi Zhang, Zhuo Chen, Lingbing Guo, Yajing Xu 等SIGIR 2024 · 被引用 39 次
- RTQ: Rethinking Video-language Understanding Based on Image-text ModelXiao Wang, Yaoyu Li, Tian Gan, Zheng Zhang 等ACM MM 2023 · 被引用 14 次
它引用的顶会 Paper20
- Frozen in Time: A Joint Video and Image Encoder for End-to-End RetrievalMax Bain, Arsha Nagrani, Gül Varol, Andrew ZissermanICCV 2021 · 被引用 1,550 次
- HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video ClipsAntoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi 等ICCV 2019 · 被引用 1,437 次
- VideoBERT: A Joint Model for Video and Language Representation LearningChen Sun, Austin Myers, Carl Vondrick, Kevin Murphy 等ICCV 2019 · 被引用 1,396 次
- Unified Vision-Language Pre-Training for Image Captioning and VQALuowei Zhou, Hamid Palangi, Lei Zhang, Houdong Hu 等AAAI 2020 · 被引用 1,047 次
- SwinBERT: End-to-End Transformers with Sparse Attention for Video CaptioningKevin Lin, Linjie Li, Chung-Ching Lin, Faisal Ahmed 等CVPR 2022 · 被引用 263 次
相关 Paper
- Cross-modal Fusion Transformer for Integrating Retrieved Knowledge into Video Caption GenerationKarina Abubakirova, Waseem Ullah, Latif U. Khan, Mohsen GuizaniKDD 2026
- Discriminative Latent Semantic Graph for Video CaptioningYang Bai, Junyan Wang, Yang Long, Bingzhang Hu 等ACM MM 2021 · 被引用 26 次
- Object Relational Graph With Teacher-Recommended Learning for Video CaptioningZiqi Zhang, Yaya Shi, Chunfeng Yuan, Bing Li 等CVPR 2020
- Leveraging Weighted Cross-Graph Attention for Visual and Semantic Enhanced Video Captioning NetworkDeepali Verma, Arya Haldar, Tanima DuttaAAAI 2023 · 被引用 13 次
- VLTinT: Visual-Linguistic Transformer-in-Transformer for Coherent Video Paragraph CaptioningKashu Yamazaki, Khoa Vo, Quang Sang Truong, Bhiksha Raj 等AAAI 2023 · 被引用 44 次
