Progressive Graph Attention Network for Video Question Answering
Liang Peng, Shuangji Yang, Yi Bin, Guoqing Wang
摘要
Video question answering (Video-QA) is a task of answering a natural language question related to the content of a video. Existing methods generally explore the single interactions between objects or between frames, which are insufficient to deal with the sophisticated scenes in videos. To tackle this problem, we propose a novel model, termed Progressive Graph Attention Network (PGAT), which can jointly explore the multiple visual relations on object-level, frame-level and clip-level. Specifically, in the object-level relation encoding, we design two kinds of complementary graphs, one for learning the spatial and semantic relations between objects from the same frame, the other for modeling the temporal relations between the same object from different frames. The frame-level graph explores the interactions between diverse frames to record the fine-grained appearance change, while the clip-level graph models the temporal and semantic relations between various actions from clips. These different-level graphs are concatenated in a progressive manner to learn the visual relations from low-level to high-level. Furthermore, we for the first time identified that there are serious answer biases with TGIF-QA, a very large Video-QA dataset, and reconstructed a new dataset based on it to overcome the biases, called TGIF-QA-R. We evaluate the proposed model on three benchmark datasets and the new TGIF-QA-R, and the experimental results demonstrate that our model significantly outperforms other state-of-the-art models. Our codes and dataset are available at https://github.com/PengLiang-cn/PGAT.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper13
- Video Question Answering: Datasets, Algorithms and ChallengesYaoyao Zhong, Wei Ji, Junbin Xiao, Yicong Li 等EMNLP 2022 · 被引用 70 次
- Your Negative May not Be True Negative: Boosting Image-Text Matching with False Negative EliminationHaoxuan Li, Yi Bin, Junrong Liao, Yang Yang 等ACM MM 2023 · 被引用 42 次
- Unifying Two-Stream Encoders with Transformers for Cross-Modal RetrievalYi Bin, Haoxuan Li, Yahui Xu, Xing Xu 等ACM MM 2023 · 被引用 33 次
- Equivariant and Invariant Grounding for Video Question AnsweringYicong Li, Xiang Wang, Junbin Xiao, Tat-Seng ChuaACM MM 2022 · 被引用 33 次
- Discovering Spatio-Temporal Rationales for Video Question AnsweringYicong Li, Junbin Xiao, Chun Feng, Xiang Wang 等ICCV 2023 · 被引用 26 次
相关 Paper
- Divide and Conquer: Question-Guided Spatio-Temporal Contextual Attention for Video Question AnsweringJianwen Jiang, Ziqiang Chen, Haojie Lin, Xibin Zhao 等AAAI 2020 · 被引用 129 次
- Location-Aware Graph Convolutional Networks for Video Question AnsweringDeng Huang, Peihao Chen, Runhao Zeng, Qing Du 等AAAI 2020 · 被引用 187 次
- Pairwise VLAD Interaction Network for Video Question AnsweringHui Wang, Dan Guo, Xian-Sheng Hua, Meng WangACM MM 2021 · 被引用 15 次
- Language-Guided Visual Aggregation Network for Video Question AnsweringXiao Liang, Di Wang, Quan Wang, Bo Wan 等ACM MM 2023 · 被引用 5 次
- Multi-Question Learning for Visual Question AnsweringChenyi Lei, Lei Wu, Dong Liu, Zhao Li 等AAAI 2020 · 被引用 9 次
