Causal-CoG: A Causal-Effect Look at Context Generation for Boosting Multi-Modal Language Models
Shitian Zhao, Zhuowan Li, Yadong Lu, Alan L. Yuille, Yan Wang
Abstract
While Multi-modal Language Models (MLMs) demonstrate impressive multimodal ability, they still struggle on providing factual and precise responses for tasks like visual question answering (VQA). In this paper, we address this challenge from the perspective of contextual information. We propose Causal Context Generation, Causal-CoG, which is a prompting strategy that engages contextual information to enhance precise VQA during inference. Specifically, we prompt MLMs to generate contexts, i.e, text description of an image, and engage the generated contexts for question answering. Moreover, we investigate the ad-vantage of contexts on VQA from a causality perspective, introducing causality filtering to select samples for which contextual information is helpful. To show the effectiveness of Causal-CoG, we run extensive experiments on 10 multimodal benchmarks and show consistent improvements, e.g., +6.30% on POPE, +13.69% on Vizwiz and +6.43% on VQAv2 compared to direct decoding, surpassing existing methods. We hope Casual-CoG inspires explorations of context knowledge in multimodal models, and serves as a plug-and-play strategy for MLM decoding.<sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup><sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup>Code is released zhaoshitian/Causal-CoG
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 244d0aa7-1294-4b1d-ad9f-877c0cff0702Cited by top-tier papers2
- Beyond Perceptual Shortcuts: Causal-Inspired Debiasing Optimization for Generalizable Video Reasoning in Lightweight MLLMsJingze Wu, Quan Zhang, Hongfei Suo, Zeqiang Cai et al.CVPR 2026 · 2 citations
- CoE: Chain-of-Explanation via Automatic Visual Concept Circuit Description and Polysemanticity QuantificationWenlong Yu, Qilong Wang, Chuang Liu, Dong Li et al.CVPR 2025
Builds on22
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo et al.NeurIPS 2022 · 8,168 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
Related papers
- QG-CoC: Question-Guided Chain-of-Captions for Large Multimodal ModelsKuei-Chun Kao, Hsu Tzu-Yin, Yunqi Hong, Ruochen Wang et al.EMNLP 2025
- Seeing Far and Clearly: Mitigating Hallucinations in MLLMs with Attention Causal DecodingFeilong Tang, Chengzhi Liu, Zhongxing Xu, Ming Hu et al.CVPR 2025
- Modality-Aware Integration with Large Language Models for Knowledge-Based Visual Question AnsweringJunnan Dong, Qinggang Zhang, Huachi Zhou, Daochen Zha et al.ACL 2024 · 11 citations
- How to Configure Good In-Context Sequence for Visual Question AnsweringLi Li, Jiawei Peng, Huiyi Chen, Chongyang Gao et al.CVPR 2024
- Debiasing Multimodal Large Language Models via Penalization of Language PriorsYifan Zhang, Yang Shi, Weichen Yu, Qingsong Wen et al.ACM MM 2025 · 6 citations
