Discovering Spatio-Temporal Rationales for Video Question Answering
Yicong Li, Junbin Xiao, Chun Feng, Xiang Wang, Tat-Seng Chua
摘要
This paper strives to solve complex video question answering (VideoQA) which features long video containing multiple objects and events at different time. To tackle the challenge, we highlight the importance of identifying question-critical temporal moments and spatial objects from the vast amount of video content. Towards this, we propose a Spatio-Temporal Rationalization (STR), a differentiable selection module that adaptively collects questioncritical moments and objects using cross-modal interaction. The discovered video moments and objects are then served as grounded rationales to support answer reasoning. Based on STR, we further propose TranSTR, a Transformerstyle neural network architecture that takes STR as the core and additionally underscores a novel answer interaction mechanism to coordinate STR for answer decoding. Experiments on four datasets show that TranSTR achieves new state-of-the-art (SoTA). Especially, on NExT-QA and Causal-VidQA which feature complex VideoQA, it significantly surpasses the previous SoTA by 5.8% and 6.8%, respectively. We then conduct extensive studies to verify the importance of STR as well as the proposed answer interaction mechanism. With the success of TranSTR and our comprehensive analysis, we hope this work can spark more future efforts in complex VideoQA. Code will be released at https://github.com/yl3800/TranSTR . Answ A rid B loo C pu D rep E tur Question: What does the second person do after stopping her bike? Answer Candidates: A ride the bike B look at the girl C push the bike D repair road E turn around Answer: push the bike stopping Top Acc (%) 100% 70% 40% 10% Top % Video Length Short Long 272649 Acc (%) 61 +0.7 -1.5 -2.0 -1.3 (a) A example of long video (52s) with multiple objects, the question-related frames and objects are located to support the reasoning. Que seco stop Answer Candida A ride the bike B look at the gir C push the bike D repair road E turn around Question: What does the second person do after stopping her bike? Answer Candidates: A ride the bike B look at the girl C push the bike D repair road E turn around Answer: push the bike stopping Top percentag Acc (%) 100% 70% 40% 10% Top Percentage by Video Length 2726497009 Acc (%) 61 +0.7 -1.5 -2.0 -1.3 (b) Accuracy by video length. Q se st Answer Cand A ride the bik B look at the C push the bik D repair road E turn around Question: What does the second person do after stopping her bike? Answer Candidates: A ride the bike B look at the girl C push the bike D repair road E turn around Answer: push the bike stopping Top percent Acc (%) 100% 70% 40% 10% Top % Video Length Short Long
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- Video-of-Thought: Step-by-Step Video Reasoning from Perception to CognitionHao Fei, Shengqiong Wu, Wei Ji, Hanwang Zhang 等ICML 2024 · 被引用 182 次
- Can I Trust Your Answer? Visually Grounded Video Question AnsweringJunbin Xiao, Angela Yao, Yicong Li, Tat-Seng ChuaCVPR 2024 · 被引用 44 次
- TimeSearch-R: Adaptive Temporal Search for Long-Form Video Understanding via Self-Verification Reinforcement LearningJunwen Pan, Qizhe Zhang, Rui Zhang, Ming Lu 等ICLR 2026 · 被引用 20 次
- Weakly Supervised Gaussian Contrastive Grounding with Large Multimodal Models for Video Question AnsweringHaibo Wang, Chenghang Lai, Yixuan Sun, Weifeng GeACM MM 2024 · 被引用 12 次
- Geometric Alignment and Prior Modulation for View-Guided Point Cloud Completion on Unseen CategoriesJingqiao Xiu, Yicong Li, Na Zhao, Han Fang 等ICCV 2025 · 被引用 2 次
它引用的顶会 Paper14
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Deberta: decoding-Enhanced Bert with Disentangled AttentionPengcheng He, Xiaodong Liu, Jianfeng Gao, Weizhu ChenICLR 2021 · 被引用 3,729 次
- MERLOT: Multimodal Neural Script Knowledge ModelsRowan Zellers, Ximing Lu, Jack Hessel, Youngjae Yu 等NeurIPS 2021 · 被引用 463 次
- Just Ask: Learning to Answer Questions from Millions of Narrated VideosAntoine Yang, Antoine Miech, Josef Sivic, Ivan Laptev 等ICCV 2021 · 被引用 345 次
- Reasoning with Heterogeneous Graph Alignment for Video Question AnsweringPin Jiang, Yahong HanAAAI 2020 · 被引用 214 次
相关 Paper
- Visual Causal Scene Refinement for Video Question AnsweringYushen Wei, Yang Liu, Hong Yan, Guanbin Li 等ACM MM 2023 · 被引用 31 次
- Rethinking Multi-Modal Alignment in Multi-Choice VideoQA from Feature and Sample PerspectivesShaoning Xiao, Long Chen, Kaifeng Gao, Zhao Wang 等EMNLP 2022 · 被引用 5 次
- Dynamic Spatio-Temporal Modular Network for Video Question AnsweringZi Qian, Xin Wang, Xuguang Duan, Hong Chen 等ACM MM 2022 · 被引用 14 次
- MIST : Multi-modal Iterative Spatial-Temporal Transformer for Long-form Video Question AnsweringDifei Gao, Luowei Zhou, Lei Ji, Linchao Zhu 等CVPR 2023
- Video Question Answering: Datasets, Algorithms and ChallengesYaoyao Zhong, Wei Ji, Junbin Xiao, Yicong Li 等EMNLP 2022 · 被引用 70 次
