Location-Aware Graph Convolutional Networks for Video Question Answering
Deng Huang, Peihao Chen, Runhao Zeng, Qing Du, Mingkui Tan, Chuang Gan
摘要
We addressed the challenging task of video question answering, which requires machines to answer questions about videos in a natural language form. Previous state-of-the-art methods attempt to apply spatio-temporal attention mechanism on video frame features without explicitly modeling the location and relations among object interaction occurred in videos. However, the relations between object interaction and their location information are very critical for both action recognition and question reasoning. In this work, we propose to represent the contents in the video as a locationaware graph by incorporating the location information of an object into the graph construction. Here, each node is associated with an object represented by its appearance and location features. Based on the constructed graph, we propose to use graph convolution to infer both the category and temporal locations of an action. As the graph is built on objects, our method is able to focus on the foreground action contents for better video question answering. Lastly, we leverage an attention mechanism to combine the output of graph convolution and encoded question features for final answer reasoning. Extensive experiments demonstrate the effectiveness of the proposed methods. Specifically, our method significantly outperforms state-of-the-art methods on TGIF-QA, Youtube2Text-QA and MSVD-QA datasets. Code and pre-trained models are publicly available at: https://github.com/SunDoge/L-GCN
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper51
- Just Ask: Learning to Answer Questions from Millions of Narrated VideosAntoine Yang, Antoine Miech, Josef Sivic, Ivan Laptev 等ICCV 2021 · 被引用 345 次
- Zero-Shot Video Question Answering via Frozen Bidirectional Language ModelsAntoine Yang, Antoine Miech, Josef Sivic, Ivan Laptev 等NeurIPS 2022 · 被引用 305 次
- Video as Conditional Graph Hierarchy for Multi-Granular Question AnsweringJunbin Xiao, Angela Yao, Zhiyuan Liu, Yicong Li 等AAAI 2022 · 被引用 145 次
- RSPNet: Relative Speed Perception for Unsupervised Video Representation LearningPeihao Chen, Deng Huang, Dongliang He, Xiang Long 等AAAI 2021 · 被引用 140 次
- Invariant Grounding for Video Question AnsweringYicong Li, Xiang Wang, Junbin Xiao, Wei Ji 等CVPR 2022 · 被引用 108 次
它引用的顶会 Paper2
- Graph Convolutional Networks for Temporal Action LocalizationRunhao Zeng, Wenbing Huang, Chuang Gan, Mingkui Tan 等ICCV 2019 · 被引用 536 次
- Multi-Agent Reinforcement Learning Based Frame Sampling for Effective Untrimmed Video RecognitionWenhao Wu, Dongliang He, Xiao Tan, Shifeng Chen 等ICCV 2019 · 被引用 135 次
相关 Paper
- Divide and Conquer: Question-Guided Spatio-Temporal Contextual Attention for Video Question AnsweringJianwen Jiang, Ziqiang Chen, Haojie Lin, Xibin Zhao 等AAAI 2020 · 被引用 129 次
- Progressive Graph Attention Network for Video Question AnsweringLiang Peng, Shuangji Yang, Yi Bin, Guoqing WangACM MM 2021 · 被引用 47 次
- Language-Guided Visual Aggregation Network for Video Question AnsweringXiao Liang, Di Wang, Quan Wang, Bo Wan 等ACM MM 2023 · 被引用 5 次
- Aligned Dual Channel Graph Convolutional Network for Visual Question AnsweringQingbao Huang, Jielong Wei, Yi Cai, Changmeng Zheng 等ACL 2020 · 被引用 79 次
- Attend What You Need: Motion-Appearance Synergistic Networks for Video Question AnsweringAhjeong Seo, Gi-Cheon Kang, Joonhan Park, Byoung-Tak ZhangACL 2021
