Pairwise VLAD Interaction Network for Video Question Answering
Hui Wang, Dan Guo, Xian-Sheng Hua, Meng Wang
Abstract
Video Question Answering (VideoQA) is a challenging problem, as it requires a joint understanding of video and natural language question. Existing methods perform correlation learning between video and question have achieved great success. However, previous methods merely model relations between individual video frames (or clips) and words, which are not enough to correctly answer the question. From human's perspective, answering a video question should first summarize both visual and language information, and then explore their correlations for answer reasoning. In this paper, we propose a new method called Pairwise VLAD Interaction Network (PVI-Net) to address this problem. Specifically, we develop a learnable clustering-based VLAD encoder to respectively summarize video and question modalities into a small number of compact VLAD descriptors. For correlation learning, a pairwise VLAD interaction mechanism is proposed to better exploit complementary information for each pair of modality descriptors, avoiding modeling uninformative individual relations (e.g., frame-word and clip-word relations), and exploring both inter-and intra-modality relations simultaneously. Experimental results show that our approach achieves state-of-the-art performance on three VideoQA datasets: TGIF-QA, MSVD-QA, and MSRVTT-QA. Visualization results further validate the interpretability of our method.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6aea4bbc-f50b-4acc-91dd-cb2cedc858aaCited by top-tier papers2
- Dynamic Spatio-Temporal Modular Network for Video Question AnsweringZi Qian, Xin Wang, Xuguang Duan, Hong Chen et al.ACM MM 2022 · 14 citations
- Efficient End-to-End Video Question Answering with Pyramidal Multimodal TransformerMin Peng, Chongyang Wang, Yu Shi, Xiang-Dong ZhouAAAI 2023 · 13 citations
Builds on5
- Reasoning with Heterogeneous Graph Alignment for Video Question AnsweringPin Jiang, Yahong HanAAAI 2020 · 214 citations
- Location-Aware Graph Convolutional Networks for Video Question AnsweringDeng Huang, Peihao Chen, Runhao Zeng, Qing Du et al.AAAI 2020 · 187 citations
- On the General Value of Evidence, and Bilingual Scene-Text Visual Question AnsweringXinyu Wang, Yuliang Liu, Chunhua Shen, Chun Chet Ng et al.CVPR 2020
- Iterative Context-Aware Graph Inference for Visual DialogDan Guo, Hui Wang, Hanwang Zhang, Zheng-Jun Zha et al.CVPR 2020
- Hierarchical Conditional Relation Networks for Video Question AnsweringThao Minh Le, Vuong Le, Svetha Venkatesh, Truyen TranCVPR 2020
Related papers
- Progressive Graph Attention Network for Video Question AnsweringLiang Peng, Shuangji Yang, Yi Bin, Guoqing WangACM MM 2021 · 47 citations
- Language-Guided Visual Aggregation Network for Video Question AnsweringXiao Liang, Di Wang, Quan Wang, Bo Wan et al.ACM MM 2023 · 5 citations
- Divide and Conquer: Question-Guided Spatio-Temporal Contextual Attention for Video Question AnsweringJianwen Jiang, Ziqiang Chen, Haojie Lin, Xibin Zhao et al.AAAI 2020 · 129 citations
- Bridge To Answer: Structure-Aware Graph Interaction Network for Video Question AnsweringJungin Park, Jiyoung Lee, Kwanghoon SohnCVPR 2021
- Modality Shifting Attention Network for Multi-Modal Video Question AnsweringJunyeong Kim, Minuk Ma, Trung X. Pham, Kyungsu Kim et al.CVPR 2020
