CausalVLBench: Benchmarking Visual Causal Reasoning in Large Vision-Language Models
Aneesh Komanduri, Karuna Bhaila, Xintao Wu
摘要
Large language models (LLMs) have shown remarkable ability in various language tasks, especially with their emergent in-context learning capability. Extending LLMs to incorporate visual inputs, large vision-language models (LVLMs) have shown impressive performance in tasks such as recognition and visual question answering (VQA). Despite increasing interest in the utility of LLMs in causal reasoning tasks such as causal discovery and counterfactual reasoning, there has been relatively little work showcasing the abilities of LVLMs on visual causal reasoning tasks. We take this opportunity to formally introduce a comprehensive causal reasoning benchmark for multi-modal in-context learning from LVLMs. Our CausalVLBench encompasses three representative tasks: causal structure inference, intervention target prediction, and counterfactual prediction. We evaluate the ability of state-ofthe-art open-source LVLMs on our causal reasoning tasks across three causal representation learning datasets and demonstrate their fundamental strengths and weaknesses. We hope that our benchmark elucidates the drawbacks of existing vision-language models and motivates new directions and paradigms in improving the visual causal reasoning abilities of LVLMs. We make our code and data available at https: //github.com/Akomand/CausalVLBench . * Light Position: left * Shadow Length: short * Shadow Position: right 2. Desired Manipulation: We want to change the light position to 'right'. 3. Causal Reasoning: * The rule (2) says that if the light position changes, it causes both the shadow length and shadow position to change. It does NOT cause the pendulum angle to change. * Given that the light moves from left to right, the shadow position will move from right to left. * Given that the light moves from left to right, the shadow length will become long. 4. Predicted Final State: * Pendulum Angle: left (unchanged) * Light Position: right * Shadow Length: long * Shadow Position: left
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- The Abstraction Gap in Vision-Language Causal ReasoningChinh Hoang, Mohammad HasanICML 2026
- From Prompts to Tokens: Internalizing Causal Supervision in Vision-Language Model for Multi-Image Causal ReasoningHaoping Yu, Yuanxi Li, Jing MaICML 2026
它引用的顶会 Paper17
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo 等NeurIPS 2022 · 被引用 8,168 次
- The Linear Representation Hypothesis and the Geometry of Large Language ModelsKiho Park, Yo Joong Choe, Victor VeitchICML 2024 · 被引用 461 次
- What matters when building vision-language models?Hugo Laurençon, Léo Tronchon, Matthieu Cord, Victor SanhNeurIPS 2024 · 被引用 401 次
相关 Paper
- VL-ICL Bench: The Devil in the Details of Multimodal In-Context LearningYongshuo Zong, Ondrej Bohdal, Timothy M. HospedalesICLR 2025
- OmniVideoBench: Towards Audio-Visual Understanding Evaluation for Omni MLLMsCaorui Li, Yu Chen, Yiyan Ji, Jin Xu 等ICLR 2026 · 被引用 53 次
- CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language ModelsZihui Cheng, Qiguang Chen, Jin Zhang, Hao Fei 等AAAI 2025 · 被引用 36 次
- Hallusionbench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language ModelsTianrui Guan, Fuxiao Liu, Xiyang Wu, Ruiqi Xian 等CVPR 2024
- CounterBench: Evaluating and Improving Counterfactual Reasoning in Large Language ModelsYuefei Chen, Vivek K. Singh, Jing Ma, Ruixiang TangAAAI 2026 · 被引用 1 次
