Measuring Compositional Consistency for Video Question Answering
Mona Gandhi, Mustafa Omer Gul, Eva Prakash, Madeleine Grunde-McLaughlin, Ranjay Krishna, Maneesh Agrawala
摘要
Recent video question answering benchmarks indicate that state-of-the-art models struggle to answer compositional questions. However, it remains unclear which types of compositional reasoning cause models to mispredict. Furthermore, it is difficult to discern whether models arrive at answers using compositional reasoning or by leveraging data biases. In this paper, we develop a question decomposition engine that programmatically deconstructs a compositional question into a directed acyclic graph of sub-questions. The graph is designed such that each parent question is a composition of its children. We present AGQA-Decomp, a benchmark containing 2.3M question graphs, with an average of 11.49 sub-questions per graph, and 4.55M total new sub-questions. Using question graphs, we evaluate three state-of-the-art models with a suite of novel compositional consistency metrics. We find that models either cannot reason correctly through most compositions or are reliant on incorrect reasoning to reach answers, frequently contradicting themselves or achieving high accuracies when failing at intermediate reasoning steps. * Equal contribution Legend: objects relationships actions time Q. What is the first object that the person is touching after taking a picture? Q. Is a phone the first object that the person is touching after taking a picture? Q. Does a phone exist? Q. Is the person touching something? Q. Is the person taking a picture? Q. Does a person exist? Q. Is the person taking something? Q. Does a picture exist? Compositional question decomposition Q. What is the person touching after taking a picture? Q. Is a person touching something after taking a picture? Q. Does a person exist after taking a picture?
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- TIFA: Accurate and Interpretable Text-to-Image Faithfulness Evaluation with Question AnsweringYushi Hu, Benlin Liu, Jungo Kasai, Yizhong Wang 等ICCV 2023 · 被引用 400 次
- Video Question Answering: Datasets, Algorithms and ChallengesYaoyao Zhong, Wei Ji, Junbin Xiao, Yicong Li 等EMNLP 2022 · 被引用 70 次
- Encoding and Controlling Global Semantics for Long-form Video Question AnsweringThong Nguyen, Zhiyuan Hu, Xiaobao Wu, Cong-Duy Nguyen 等EMNLP 2024 · 被引用 3 次
- Counterfactual Evolution of Multimodal Datasets via Visual ProgrammingMinghe Gao, Zhongqi Yue, Wenjie Yan, Yihao Hu 等NeurIPS 2025 · 被引用 1 次
- @ CREPE: Can Vision-Language Foundation Models Reason Compositionally?Zixian Ma, Jerry Hong, Mustafa Omer Gul, Mona Gandhi 等CVPR 2023
它引用的顶会 Paper12
- CLEVRER: Collision Events for Video Representation and ReasoningKexin Yi, Chuang Gan, Yunzhu Li, Pushmeet Kohli 等ICLR 2020 · 被引用 584 次
- CATER: A diagnostic dataset for Compositional Actions & TEmporal ReasoningRohit Girdhar, Deva RamananICLR 2020 · 被引用 198 次
- Understanding and Evaluating Racial Biases in Image CaptioningDora Zhao, Angelina Wang, Olga RussakovskyICCV 2021 · 被引用 165 次
- Explaining Answers with Entailment TreesBhavana Dalvi, Peter Jansen, Oyvind Tafjord, Zhengnan Xie 等EMNLP 2021 · 被引用 6 次
- Unsupervised Question Decomposition for Question AnsweringEthan Perez, Patrick Lewis, Wen-tau Yih, Kyunghyun Cho 等EMNLP 2020 · 被引用 6 次
相关 Paper
- Align and Aggregate: Compositional Reasoning with Video Alignment and Answer Aggregation for Video Question-AnsweringZhaohe Liao, Jiangtong Li, Li Niu, Liqing ZhangCVPR 2024
- AGQA: A Benchmark for Compositional Spatio-Temporal ReasoningMadeleine Grunde-McLaughlin, Ranjay Krishna, Maneesh AgrawalaCVPR 2021
- ANetQA: A Large-scale Benchmark for Fine-grained Compositional Reasoning over Untrimmed VideosZhou Yu, Lixiang Zheng, Zhou Zhao, Fei Wu 等CVPR 2023
- Commonsense Video Question Answering through Video-Grounded Entailment Tree ReasoningHuabin Liu, Filip Ilievski, Cees G. M. SnoekCVPR 2025
- Maintaining Reasoning Consistency in Compositional Visual Question AnsweringChenchen Jing, Yunde Jia, Yuwei Wu, Xinyu Liu 等CVPR 2022 · 被引用 27 次
