ACRE: Abstract Causal REasoning Beyond Covariation
Chi Zhang, Baoxiong Jia, Mark Edmonds, Song-Chun Zhu, Yixin Zhu
摘要
Causal induction, i.e., identifying unobservable mechanisms that lead to the observable relations among variables, has played a pivotal role in modern scientific discovery, especially in scenarios with only sparse and limited data. Humans, even young toddlers, can induce causal relationships surprisingly well in various settings despite its notorious difficulty. However, in contrast to the commonplace trait of human cognition is the lack of a diagnostic benchmark to measure causal induction for modern Artificial Intelligence (AI) systems. Therefore, in this work, we introduce the Abstract Causal REasoning (ACRE) dataset for systematic evaluation of current vision systems in causal induction. Motivated by the stream of research on causal discovery in Blicket experiments, we query a visual reasoning system with the following four types of questions in either an independent scenario or an interventional scenario: direct, indirect, screening-off, and backward-blocking, intentionally going beyond the simple strategy of inducing causal relationships by covariation. By analyzing visual reasoning architectures on this testbed, we notice that pure neural models tend towards an associative strategy under their chance-level performance, whereas neuro-symbolic combinations struggle in backward-blocking reasoning. These deficiencies call for future research in models with a more comprehensive capability of causal induction.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper21
- Phenomenal Yet Puzzling: Testing Inductive Reasoning Capabilities of Language Models with Hypothesis RefinementLinlu Qiu, Liwei Jiang, Ximing Lu, Melanie Sclar 等ICLR 2024 · 被引用 114 次
- Attention over Learned Object Embeddings Enables Complex Visual ReasoningDavid Ding, Felix Hill, Adam Santoro, Malcolm Reynolds 等NeurIPS 2021 · 被引用 87 次
- MR-Ben: A Meta-Reasoning Benchmark for Evaluating System-2 Thinking in LLMsZhongshen Zeng, Yinhong Liu, Yingjia Wan, Jingyao Li 等NeurIPS 2024 · 被引用 51 次
- Doing Experiments and Revising Rules with Natural Language and Probabilistic ReasoningTop Piriyakulkij, Cassidy Langenfeld, Tuan Anh Le, Kevin EllisNeurIPS 2024 · 被引用 30 次
- MEWL: Few-shot multimodal word learning with referential uncertaintyGuangyuan Jiang, Manjie Xu, Shiji Xin, Wei Liang 等ICML 2023 · 被引用 29 次
它引用的顶会 Paper10
- CLEVRER: Collision Events for Video Representation and ReasoningKexin Yi, Chuang Gan, Yunzhu Li, Pushmeet Kohli 等ICLR 2020 · 被引用 584 次
- CATER: A diagnostic dataset for Compositional Actions & TEmporal ReasoningRohit Girdhar, Deva RamananICLR 2020 · 被引用 198 次
- A Benchmark for Systematic Generalization in Grounded Language UnderstandingLaura Ruis, Jacob Andreas, Marco Baroni, Diane Bouchacourt 等NeurIPS 2020 · 被引用 169 次
- Permutation Equivariant Models for Compositional Generalization in LanguageJonathan Gordon, David Lopez-Paz, Marco Baroni, Diane BouchacourtICLR 2020 · 被引用 112 次
- CoPhy: Counterfactual Learning of Physical DynamicsFabien Baradel, Natalia Neverova, Julien Mille, Greg Mori 等ICLR 2020 · 被引用 105 次
相关 Paper
- Visual Abductive ReasoningChen Liang, Wenguan Wang, Tianfei Zhou, Yi YangCVPR 2022 · 被引用 50 次
- CELLO: Causal Evaluation of Large Vision-Language ModelsMeiqi Chen, Bo Peng, Yan Zhang, Chaochao LuEMNLP 2024 · 被引用 4 次
- Your Reasoning Benchmark May Not Test Reasoning: Revealing Perception Bottleneck in Abstract Reasoning BenchmarksXinhe Wang, Jin Huang, Xingjian Zhang, Tianhao Wang 等ACL 2026 · 被引用 3 次
- QUITE: Quantifying Uncertainty in Natural Language Text in Bayesian Reasoning ScenariosTimo Pierre Schrader, Lukas Lange, Simon Razniewski, Annemarie FriedrichEMNLP 2024
- Zebra-CoT: A Dataset for Interleaved Vision-Language ReasoningAng Li, Charles L. Wang, Deqing Fu, Kaiyu Yue 等ICLR 2026 · 被引用 85 次
