Inferential Visual Question Generation
Chao Bi, Shuhui Wang, Zhe Xue, Shengbo Chen, Qingming Huang
Abstract
The task of Visual Question Generation (VQG) aims to generate natural language questions for images. Many methods regard it as a reverse Visual Question Answering (VQA) task. They trained a data-driven generator on VQA datasets, which is hard to obtain questions that can challenge robots and humans. Other methods rely heavily on elaborate but expensive artificial preprocessing to generate. To overcome these limitations, we propose a method to generate inferential questions from the image with noisy captions. Our method first introduces a core scene graph generation module, which can align text features and salient visual features to the initial scene graph. It constructs a special core scene graph with expanded linkage outwards from the high-confidence nodes hop by hop. Next, a question generation module uses the core scene graph as a basis to instantiate the function templates, resulting in questions with varying inferential paths. Experiments show that the visual questions generated by our method are controllable in both content and difficulty, and demonstrate clear inferential properties. In addition, since the salient region, captions, and function templates can be replaced by human-customized ones, our method has strong scalability and potential for more interactive applications. Finally, we use our method to automatically build a new dataset, InVQA, containing about 120k images and 480k question-answer pairs, to facilitate the development of more versatile VQA models.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get d4e33ff0-8680-4c4a-ae0b-a04a4657895dCited by top-tier papers1
Ask how each one uses itRelated papers
- Deconfounded Visual Question Generation with Causal InferenceJiali Chen, Zhenjun Guo, Jiayuan Xie, Yi Cai et al.ACM MM 2023 · 8 citations
- VrR-VG: Refocusing Visually-Relevant RelationshipsYuanzhi Liang, Yalong Bai, Wei Zhang, Xueming Qian et al.ICCV 2019 · 93 citations
- Learning to Generate Visual Questions with Noisy SupervisionKai Shen, Lingfei Wu, Siliang Tang, Yueting Zhuang et al.NeurIPS 2021 · 13 citations
- Explicitly Guided Difficulty-Controllable Visual Question GenerationJiayuan Xie, Mengqiu Cheng, Xinting Zhang, Yi Cai et al.AAAI 2025 · 2 citations
- Multiple Objects-Aware Visual Question GenerationJiayuan Xie, Yi Cai, Qingbao Huang, Tao WangACM MM 2021 · 23 citations
