From Prompts to Tokens: Internalizing Causal Supervision in Vision-Language Model for Multi-Image Causal Reasoning
Haoping Yu, Yuanxi Li, Jing Ma
Abstract
Visual causal reasoning is essential for understanding and intervening in the physical world, requiring identification of causal variables from visual inputs and reasoning over intervention effects. Despite recent progress, large vision-language models (VLMs) remain brittle at such tasks, especially for interventional and counterfactual queries over multi-image inputs. Most existing explorations inject causal knowledge via textual prompts, leaving causal mechanisms external to model execution and limiting reliable control during inference. To address this problem, we propose BridgeVLM, which internalizes visual causal reasoning by inducing a causal graph from multi-image inputs and converting it into structured Causal Tokens executed by RAMP layers injected into the LLM decoder for causal message passing. We further introduce a unified training interface M3S for fine-grained causal supervision from different granularities (local/global level). BridgeVLM achieves 54.4% accuracy on intervention tasks on CausalVLBench (vs. 33.2% with prompt-level supervision), improves results on Causal3D from 43.6% to 49.0%, and substantially improves causal structure learning on CausalVLBench (: 33.4% 75.1%).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on10
- Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought PromptingMiles Turpin, Julian Michael, Ethan Perez, Samuel R. BowmanNeurIPS 2023 · 1,792 citations
- On the Connection Between MPNN and Graph TransformerChen Cai, Truong Son Hy, Rose Yu, Yusu WangICML 2023 · 82 citations
- Vision Language Models are BiasedAn Vo, Khai-Nguyen Nguyen, Mohammad Reza Taesiri, Thi Tuong Vy Dang et al.ICLR 2026 · 68 citations
- Causal Prompting: Debiasing Large Language Model Prompting Based on Front-Door AdjustmentCongzhi Zhang, Linhai Zhang, Jialong Wu, Yulan He et al.AAAI 2025 · 42 citations
- Rethinking Misalignment in Vision-Language Model Adaptation from a Causal PerspectiveYanan Zhang, Jiangmeng Li, Lixiang Liu, Wenwen QiangNeurIPS 2024 · 16 citations
Related papers
- CausalVLBench: Benchmarking Visual Causal Reasoning in Large Vision-Language ModelsAneesh Komanduri, Karuna Bhaila, Xintao WuEMNLP 2025
- CF-VLM: CounterFactual Vision-Language Fine-tuningJusheng Zhang, Kaitong Cai, Yijia Fan, Jian Wang et al.NeurIPS 2025 · 71 citations
- Unveiling the Compositional Ability Gap in Vision-Language Reasoning ModelTianle Li, Jihai Zhang, Yongming Rao, Yu ChengNeurIPS 2025 · 17 citations
- Linear Mechanisms for Spatiotemporal Reasoning in Vision Language ModelsRaphaela Kang, Hongqiao Chen, Georgia Gkioxari, Pietro PeronaICLR 2026 · 11 citations
- See It, Say It, Sorted: An Iterative Training-Free Framework for Visually-Grounded Multimodal Reasoning in LVLMsYongchang Zhang, Xianzheng Ma, Tianyi Liu, Guangquan Zhou et al.CVPR 2026 · 2 citations
