Deconfounded Visual Question Generation with Causal Inference
Jiali Chen, Zhenjun Guo, Jiayuan Xie, Yi Cai, Qing Li
Abstract
Visual Question Generation (VQG) task aims to generate meaningful and logically reasonable questions about the given image targeting an answer. Existing methods mainly focus on the visual concepts present in the image for question generation and have shown remarkable performance in VQG. However, these models frequently learn highly co-occurring object relationships and attributes, which is an inherent bias in question generation. This previously overlooked bias causes models to over-exploit the spurious correlations among visual features, the target answer, and the question. Therefore, they may generate inappropriate questions that contradict the visual content or facts. In this paper, we first introduce a causal perspective on VQG and adopt the causal graph to analyze spurious correlations among variables. Building on the analysis, we propose a Knowledge Enhanced Causal Visual Question Generation (KECVQG) model to mitigate the impact of spurious correlations in question generation. Specifically, an interventional visual feature extractor (IVE) is introduced in KECVQG, which aims to obtain unbiased visual features by disentangling. Then a knowledge-guided representation extractor (KRE) is employed to align unbiased features with external knowledge. Finally, the output features from KRE are sent into a standard transformer decoder to generate questions. Extensive experiments on the VQA v2.0 and OKVQA datasets show that KECVQG significantly outperforms existing models.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get d0b4684e-bb28-4fd4-bba5-50f24c064288Cited by top-tier papers5
- Few-Shot Joint Multimodal Entity-Relation Extraction via Knowledge-Enhanced Cross-modal Prompt ModelLi Yuan, Yi Cai, Junsheng HuangACM MM 2024 · 9 citations
- EduDiag: A Benchmark for Educational Diagnostic Reasoning with Error Tracing and Correction on Large Multimodal ModelsJiali Chen, Yuqi Xue, Xusen Hei, DingBa Fu et al.CVPR 2026 · 1 citation
- AdaCoder: Adaptive Prompt Compression for Programmatic Visual Question AnsweringMahiro Ukai, Shuhei Kurita, Atsushi Hashimoto, Yoshitaka Ushiku et al.ACM MM 2024 · 1 citation
- ExpStar: Towards Automatic Commentary Generation for Multi-discipline Scientific ExperimentsJiali Chen, Yujie Jia, Zihan Wu, Jinyu Yang et al.ACM MM 2025 · 1 citation
- Diagram-Driven Course Questions GenerationXinyu Zhang, Lingling Zhang, Yanrui Wu, Muye Huang et al.EMNLP 2025
Related papers
- Inferential Visual Question GenerationChao Bi, Shuhui Wang, Zhe Xue, Shengbo Chen et al.ACM MM 2022 · 8 citations
- Multiple Objects-Aware Visual Question GenerationJiayuan Xie, Yi Cai, Qingbao Huang, Tao WangACM MM 2021 · 23 citations
- Variational Causal Inference Network for Explanatory Visual Question AnsweringDizhan Xue, Shengsheng Qian, Changsheng XuICCV 2023 · 19 citations
- ConVQG: Contrastive Visual Question Generation with Multimodal GuidanceLi Mi, Syrielle Montariol, Javiera Castillo Navarro, Xianjie Dai et al.AAAI 2024 · 8 citations
- Generative Bias for Robust Visual Question AnsweringJae-Won Cho, Dong-Jin Kim, Hyeonggon Ryu, In So KweonCVPR 2023
