Invariant Grounding for Video Question Answering
Yicong Li, Xiang Wang, Junbin Xiao, Wei Ji, Tat-Seng Chua
Abstract
Video Question Answering (VideoQA) is the task of an-swering questions about a video. At its core is understanding the alignments between visual scenes in video and linguistic semantics in question to yield the answer. In leading VideoQA models, the typical learning objective, empirical risk minimization (ERM), latches on superficial correlations between video-question pairs and answers as the alignments. However, ERM can be problematic, because it tends to over-exploit the spurious correlations between question-irrelevant scenes and answers, instead of inspecting the causal effect of question-critical scenes. As a result, the VideoQA models suffer from unreliable reasoning. In this work, we first take a causal look at VideoQA and argue that invariant grounding is the key to ruling out the spurious correlations. Towards this end, we propose a new learning framework, Invariant Grounding for VideoQA (IGV), to ground the question-critical scene, whose causal relations with answers are invariant across different interventions on the complement. With IGV, the VideoQA mod-els are forced to shield the answering process from the negative influence of spurious correlations, which significantly improves the reasoning ability. Experiments on three benchmark datasets validate the superiority of IGV in terms of accuracy, visual explainability, and generalization ability over the leading baselines. Our code is available at https://github.com/y13800/IGV.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a756c691-ca38-4d16-9bbb-f3f839ee88d3Cited by top-tier papers54
- NExT-GPT: Any-to-Any Multimodal LLMShengqiong Wu, Hao Fei, Leigang Qu, Wei Ji et al.ICML 2024 · 786 citations
- Let Invariant Rationale Discovery Inspire Graph Contrastive LearningSihang Li, Xiang Wang, An Zhang, Yingxin Wu et al.ICML 2022 · 117 citations
- Vitron: A Unified Pixel-level Vision LLM for Understanding, Generating, Segmenting, EditingHao Fei, Shengqiong Wu, Hanwang Zhang, Tat-Seng Chua et al.NeurIPS 2024 · 100 citations
- Video Question Answering: Datasets, Algorithms and ChallengesYaoyao Zhong, Wei Ji, Junbin Xiao, Yicong Li et al.EMNLP 2022 · 70 citations
- Counterfactual Reasoning for Out-of-distribution Multimodal Sentiment AnalysisTeng Sun, Wenjie Wang, Liqiang Jing, Yiran Cui et al.ACM MM 2022 · 65 citations
Builds on11
- Out-of-Distribution Generalization via Risk Extrapolation (REx)David Krueger, Ethan Caballero, Jörn-Henrik Jacobsen, Amy Zhang et al.ICML 2021 · 1,163 citations
- Environment Inference for Invariant LearningElliot Creager, Jörn-Henrik Jacobsen, Richard S. ZemelICML 2021 · 454 citations
- Reasoning with Heterogeneous Graph Alignment for Video Question AnsweringPin Jiang, Yahong HanAAAI 2020 · 214 citations
- Location-Aware Graph Convolutional Networks for Video Question AnsweringDeng Huang, Peihao Chen, Runhao Zeng, Qing Du et al.AAAI 2020 · 187 citations
- Causal Attention for Unbiased Visual RecognitionTan Wang, Chang Zhou, Qianru Sun, Hanwang ZhangICCV 2021 · 162 citations
Related papers
- Equivariant and Invariant Grounding for Video Question AnsweringYicong Li, Xiang Wang, Junbin Xiao, Tat-Seng ChuaACM MM 2022 · 33 citations
- Visual Causal Scene Refinement for Video Question AnsweringYushen Wei, Yang Liu, Hong Yan, Guanbin Li et al.ACM MM 2023 · 31 citations
- Discovering the Real Association: Multimodal Causal Reasoning in Video Question AnsweringChuanqi Zang, Hanqing Wang, Mingtao Pei, Wei LiangCVPR 2023
- Cross-modal Causal Relation Alignment for Video Question GroundingWeixing Chen, Yang Liu, Binglin Chen, Jiandong Su et al.CVPR 2025
- Can I Trust Your Answer? Visually Grounded Video Question AnsweringJunbin Xiao, Angela Yao, Yicong Li, Tat-Seng ChuaCVPR 2024 · 44 citations
