The Abstraction Gap in Vision-Language Causal Reasoning
Chinh Hoang, Mohammad Hasan
Abstract
Vision-language models (VLMs) generate fluent causal explanations, but current evaluations cannot distinguish linguistic plausibility from faithful causal reasoning. We introduce a dual-probe methodology that isolates these properties. The Text-Only Probe measures linguistic quality. The Chain-Text Probe requires models to first generate explicit causal chains. The Abstraction Gap (AG) metric quantifies the normalized performance difference. Evaluating eight VLMs on CAGE (Causal Abstraction Gap Evaluation), a benchmark of 49,500 questions across 5,500 images spanning Pearl's causal hierarchy, we find seven models exhibit AG exceeding 0.50 with text scores of 6--8 but chain scores below 2.5. Fine-tuning on 45,000 chain-annotated examples fails to close the gap. However, one model achieves near-zero AG. The capability exists within current VLM architectures and depends on pretraining and architectural choices. CAGE provides a diagnostic tool for assessing faithful causal reasoning in VLMs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6e0b3a50-9c3f-4578-a226-e36bd9f4da12Builds on29
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo et al.NeurIPS 2022 · 8,168 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
Related papers
- CaST-Bench: Benchmarking Causal Chain-Grounded Spatio-Temporal Reasoning for Video Question AnsweringMingfang Zhang, Jingjing Pan, Ashutosh Kumar, Rajat Saini et al.CVPR 2026 · 1 citation
- Reasoning Elicitation in Language Models via Counterfactual FeedbackAlihan Hüyük, Xinnuo Xu, Jacqueline R. M. A. Maasch, Aditya V. Nori et al.ICLR 2025
- CausalStep: A Benchmark for Explicit Stepwise Causal Reasoning in VideosXuchen Li, Xuzhao Li, Shiyu Hu, Kaiqi Huang et al.AAAI 2026 · 7 citations
- NoisyCausal: A Benchmark for Evaluating Causal Reasoning Under Structured NoiseZhi Xu, Yun FuACL 2026
- A Causal Lens for Evaluating Faithfulness MetricsKerem Zaman, Shashank SrivastavaEMNLP 2025
