AGQA: A Benchmark for Compositional Spatio-Temporal Reasoning
Madeleine Grunde-McLaughlin, Ranjay Krishna, Maneesh Agrawala
Abstract
Visual events are a composition of temporal actions involving actors spatially interacting with objects. When developing computer vision models that can reason about compositional spatio-temporal events, we need benchmarks that can analyze progress and uncover shortcomings. Existing video question answering benchmarks are useful, but they often conflate multiple sources of error into one accuracy metric and have strong biases that models can exploit, making it difficult to pinpoint model weaknesses. We present Action Genome Question Answering (AGQA), a new benchmark for compositional spatio-temporal reasoning. AGQA contains 192M unbalanced question answer pairs for 9.6K videos. We also provide a balanced subset of 3.9M question answer pairs, 3 orders of magnitude larger than existing benchmarks, that minimizes bias by balancing the answer distributions and types of question structures. Although human evaluators marked 86.02% of our question-answer pairs as correct, the best model achieves only 47.74% accuracy. In addition, AGQA introduces multiple training/test splits to test for various reasoning abilities, including generalization to novel compositions, to indirect references, and to more compositional steps. Using AGQA, we evaluate modern visual reasoning systems, demonstrating that the best models barely perform better than non-visual baselines exploiting linguistic biases and that none of the existing models generalize to novel compositions unseen during training. Q: What did the person hold after putting a phone somewhere? Q: Were they taking a picture or holding a bottle for longer? Q: Did they take a picture before or after they did the longest action? G99VH.mp4 Template: What did they <relation> last <time><action>, a <object1> or <object2> ?
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6ce2a76e-e657-4bc4-ac70-e5d09d3ae103Cited by top-tier papers61
- TIFA: Accurate and Interpretable Text-to-Image Faithfulness Evaluation with Question AnsweringYushi Hu, Benlin Liu, Jungo Kasai, Yizhong Wang et al.ICCV 2023 · 400 citations
- Video Question Answering: Datasets, Algorithms and ChallengesYaoyao Zhong, Wei Ji, Junbin Xiao, Yicong Li et al.EMNLP 2022 · 70 citations
- ComPhy: Compositional Physical Reasoning of Objects and Events from VideosZhenfang Chen, Kexin Yi, Yunzhu Li, Mingyu Ding et al.ICLR 2022 · 67 citations
- Compositional Temporal Grounding with Structured Variational Cross-Graph Correspondence LearningJuncheng Li, Junlin Xie, Long Qian, Linchao Zhu et al.CVPR 2022 · 63 citations
- AVQA: A Dataset for Audio-Visual Question Answering on VideosPinci Yang, Xin Wang, Xuguang Duan, Hong Chen et al.ACM MM 2022 · 60 citations
Builds on11
- CLEVRER: Collision Events for Video Representation and ReasoningKexin Yi, Chuang Gan, Yunzhu Li, Pushmeet Kohli et al.ICLR 2020 · 584 citations
- Measuring Compositional Generalization: A Comprehensive Method on Realistic DataDaniel Keysers, Nathanael Schärli, Nathan Scales, Hylke Buisman et al.ICLR 2020 · 401 citations
- Specifying Object Attributes and Relations in Interactive Scene GenerationOron Ashual, Lior WolfICCV 2019 · 190 citations
- TVQA+: Spatio-Temporal Grounding for Video Question AnsweringJie Lei, Licheng Yu, Tamara L. Berg, Mohit BansalACL 2020 · 173 citations
- Answering Questions about Charts and Generating Visual ExplanationsDae Hyun Kim, Enamul Hoque, Maneesh AgrawalaCHI 2020 · 121 citations
Related papers
- ANetQA: A Large-scale Benchmark for Fine-grained Compositional Reasoning over Untrimmed VideosZhou Yu, Lixiang Zheng, Zhou Zhao, Fei Wu et al.CVPR 2023
- Measuring Compositional Consistency for Video Question AnsweringMona Gandhi, Mustafa Omer Gul, Eva Prakash, Madeleine Grunde-McLaughlin et al.CVPR 2022 · 11 citations
- AI-VQA: Visual Question Answering based on Agent Interaction with InterpretabilityRengang Li, Cong Xu, Zhenhua Guo, Baoyu Fan et al.ACM MM 2022 · 7 citations
- Social Genome: Grounded Social Reasoning Abilities of Multimodal ModelsLeena Mathur, Marian Qian, Paul Pu Liang, Louis-Philippe MorencyEMNLP 2025
- Learning to Answer Questions in Dynamic Audio-Visual ScenariosGuangyao Li, Yake Wei, Yapeng Tian, Chenliang Xu et al.CVPR 2022 · 101 citations
