CRIPP-VQA: Counterfactual Reasoning about Implicit Physical Properties via Video Question Answering
Maitreya Patel, Tejas Gokhale, Chitta Baral, Yezhou Yang
Abstract
Videos often capture objects, their visible properties, their motion, and the interactions between different objects. Objects also have physical properties such as mass, which the imaging pipeline is unable to directly capture. However, these properties can be estimated by utilizing cues from relative object motion and the dynamics introduced by collisions. In this paper, we introduce CRIPP-VQA 1 , a new video question answering dataset for reasoning about the implicit physical properties of objects in a scene. CRIPP-VQA contains videos of objects in motion, annotated with questions that involve counterfactual reasoning about the effect of actions, questions about planning in order to reach a goal, and descriptive questions about visible properties of objects. The CRIPP-VQA test set enables evaluation under several outof-distribution settings -videos with objects with masses, coefficients of friction, and initial velocities that are not observed in the training distribution. Our experiments reveal a surprising and significant performance gap in terms of answering questions about implicit properties (the focus of this paper) and explicit properties of objects (the focus of prior work).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f34813e7-42ec-46ba-856a-d8b61f0b29e4Cited by top-tier papers7
- ContPhy: Continuum Physical Concept Learning and Reasoning from VideosZhicheng Zheng, Xin Yan, Zhenfang Chen, Jingzhou Wang et al.ICML 2024 · 22 citations
- MUST: An Effective and Scalable Framework for Multimodal Search of Target ModalityMengzhao Wang, Xiangyu Ke, Xiaoliang Xu, Lu Chen et al.ICDE 2024 · 16 citations
- DeFacto: Counterfactual Thinking with Images for Enforcing Evidence-Grounded and Faithful ReasoningTianrun Xu, Haoda Jing, Ye Li, Yuquan Wei et al.ICML 2026 · 8 citations
- SOK-Bench: A Situated Video Reasoning Benchmark with Aligned Open-World KnowledgeAndong Wang, Bo Wu, Sunli Chen, Zhenfang Chen et al.CVPR 2024 · 7 citations
- Compositional 4D Dynamic Scenes Understanding with Physics Priors for Video Question AnsweringXingrui Wang, Wufei Ma, Angtian Wang, Shuo Chen et al.ICLR 2025
Builds on14
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao et al.AAAI 2020 · 2,916 citations
- CLEVRER: Collision Events for Video Representation and ReasoningKexin Yi, Chuang Gan, Yunzhu Li, Pushmeet Kohli et al.ICLR 2020 · 584 citations
- CATER: A diagnostic dataset for Compositional Actions & TEmporal ReasoningRohit Girdhar, Deva RamananICLR 2020 · 198 citations
- Winoground: Probing Vision and Language Models for Visio-Linguistic CompositionalityTristan Thrush, Ryan Jiang, Max Bartolo, Amanpreet Singh et al.CVPR 2022 · 179 citations
- TVQA+: Spatio-Temporal Grounding for Video Question AnsweringJie Lei, Licheng Yu, Tamara L. Berg, Mohit BansalACL 2020 · 173 citations
Related papers
- ComPhy: Compositional Physical Reasoning of Objects and Events from VideosZhenfang Chen, Kexin Yi, Yunzhu Li, Mingyu Ding et al.ICLR 2022 · 67 citations
- Dynamic Visual Reasoning by Learning Differentiable Physics Models from Video and LanguageMingyu Ding, Zhenfang Chen, Tao Du, Ping Luo et al.NeurIPS 2021 · 90 citations
- Grounding Physical Concepts of Objects and Events Through Dynamic Visual ReasoningZhenfang Chen, Jiayuan Mao, Jiajun Wu, Kwan-Yee Kenneth Wong et al.ICLR 2021 · 13 citations
- ACQUIRED: A Dataset for Answering Counterfactual Questions In Real-Life VideosTe-Lin Wu, Zi-Yi Dou, Qingyuan Hu, Yu Hou et al.EMNLP 2023 · 3 citations
- QUANTIPHY: A Quantitative Benchmark Evaluating Physical Reasoning Abilities of Vision-Language ModelsLi Puyin, Tiange Xiang, Ella Mao, Shirley Wei et al.CVPR 2026 · 23 citations
