Language-Guided Visual Aggregation Network for Video Question Answering
Xiao Liang, Di Wang, Quan Wang, Bo Wan, Lingling An, Lihuo He
Abstract
Video Question Answering (VideoQA) aims to comprehend intricate relationships, actions, and events within video content, as well as the inherent links between objects and scenes, to answer text-based questions accurately. Transferring knowledge from the cross-modal pre-trained model CLIP is a natural approach, but its dual-tower structure hinders fine-grained modality interaction, posing challenges for direct application to VideoQA tasks. To address this issue, we introduce a Language-Guided Visual Aggregation (LGVA) network. It employs CLIP as an effective feature extractor to obtain language-aligned visual features with different granularities and avoids resource-intensive video pre-training. The LGVA network progressively aggregates visual information in a bottom-up manner, focusing on both regional and temporal levels, and ultimately facilitating accurate answer prediction. More specifically, it employs local cross-attention to combine pre-extracted question tokens and region embeddings, pinpointing the object of interest in the question. Then, graph attention is utilized to aggregate regions at the frame level and integrate additional captions for enhanced detail. Following this, global cross-attention is used to merge sentence and frame-level embeddings, identifying the video segment relevant to the question. Ultimately, contrastive learning is applied to optimize the similarities between aggregated visual and answer embeddings, unifying upstream and downstream tasks. Our method conserves resources by avoiding large-scale video pre-training and simultaneously demonstrates commendable performance on the NExT-QA, MSVD-QA, MSRVTT-QA, TGIF-QA, and ActivityNet-QA datasets, even outperforming some end-to-end trained models. Our code is available at https://github.com/ecoxial2007/LGVA_VideoQA.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 9bda1d9c-e65c-4a8c-bd34-6500b5350194Related papers
- Can I Trust Your Answer? Visually Grounded Video Question AnsweringJunbin Xiao, Angela Yao, Yicong Li, Tat-Seng ChuaCVPR 2024 · 44 citations
- Divide and Conquer: Question-Guided Spatio-Temporal Contextual Attention for Video Question AnsweringJianwen Jiang, Ziqiang Chen, Haojie Lin, Xibin Zhao et al.AAAI 2020 · 129 citations
- Align and Aggregate: Compositional Reasoning with Video Alignment and Answer Aggregation for Video Question-AnsweringZhaohe Liao, Jiangtong Li, Li Niu, Liqing ZhangCVPR 2024
- Progressive Graph Attention Network for Video Question AnsweringLiang Peng, Shuangji Yang, Yi Bin, Guoqing WangACM MM 2021 · 47 citations
- Long-Form Video-Language Pre-Training with Multimodal Temporal Contrastive LearningYuchong Sun, Hongwei Xue, Ruihua Song, Bei Liu et al.NeurIPS 2022 · 91 citations
