Deconfounded Visual Question Generation with Causal Inference
Jiali Chen, Zhenjun Guo, Jiayuan Xie, Yi Cai, Qing Li
摘要
Visual Question Generation (VQG) task aims to generate meaningful and logically reasonable questions about the given image targeting an answer. Existing methods mainly focus on the visual concepts present in the image for question generation and have shown remarkable performance in VQG. However, these models frequently learn highly co-occurring object relationships and attributes, which is an inherent bias in question generation. This previously overlooked bias causes models to over-exploit the spurious correlations among visual features, the target answer, and the question. Therefore, they may generate inappropriate questions that contradict the visual content or facts. In this paper, we first introduce a causal perspective on VQG and adopt the causal graph to analyze spurious correlations among variables. Building on the analysis, we propose a Knowledge Enhanced Causal Visual Question Generation (KECVQG) model to mitigate the impact of spurious correlations in question generation. Specifically, an interventional visual feature extractor (IVE) is introduced in KECVQG, which aims to obtain unbiased visual features by disentangling. Then a knowledge-guided representation extractor (KRE) is employed to align unbiased features with external knowledge. Finally, the output features from KRE are sent into a standard transformer decoder to generate questions. Extensive experiments on the VQA v2.0 and OKVQA datasets show that KECVQG significantly outperforms existing models.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper5
- Few-Shot Joint Multimodal Entity-Relation Extraction via Knowledge-Enhanced Cross-modal Prompt ModelLi Yuan, Yi Cai, Junsheng HuangACM MM 2024 · 被引用 9 次
- EduDiag: A Benchmark for Educational Diagnostic Reasoning with Error Tracing and Correction on Large Multimodal ModelsJiali Chen, Yuqi Xue, Xusen Hei, DingBa Fu 等CVPR 2026 · 被引用 1 次
- AdaCoder: Adaptive Prompt Compression for Programmatic Visual Question AnsweringMahiro Ukai, Shuhei Kurita, Atsushi Hashimoto, Yoshitaka Ushiku 等ACM MM 2024 · 被引用 1 次
- ExpStar: Towards Automatic Commentary Generation for Multi-discipline Scientific ExperimentsJiali Chen, Yujie Jia, Zihan Wu, Jinyu Yang 等ACM MM 2025 · 被引用 1 次
- Diagram-Driven Course Questions GenerationXinyu Zhang, Lingling Zhang, Yanrui Wu, Muye Huang 等EMNLP 2025
相关 Paper
- Inferential Visual Question GenerationChao Bi, Shuhui Wang, Zhe Xue, Shengbo Chen 等ACM MM 2022 · 被引用 8 次
- Multiple Objects-Aware Visual Question GenerationJiayuan Xie, Yi Cai, Qingbao Huang, Tao WangACM MM 2021 · 被引用 23 次
- Variational Causal Inference Network for Explanatory Visual Question AnsweringDizhan Xue, Shengsheng Qian, Changsheng XuICCV 2023 · 被引用 19 次
- ConVQG: Contrastive Visual Question Generation with Multimodal GuidanceLi Mi, Syrielle Montariol, Javiera Castillo Navarro, Xianjie Dai 等AAAI 2024 · 被引用 8 次
- Generative Bias for Robust Visual Question AnsweringJae-Won Cho, Dong-Jin Kim, Hyeonggon Ryu, In So KweonCVPR 2023
