Explicitly Guided Difficulty-Controllable Visual Question Generation
Jiayuan Xie, Mengqiu Cheng, Xinting Zhang, Yi Cai, Guimin Hu, Mengying Xie, Qing Li
Abstract
Visual question generation (VQG) aims to generate questions from images automatically. While existing studies primarily focus on the quality of generated questions, such as fluency and relevance, the difficulty of the questions is also a crucial factor in assessing their quality. Question difficulty directly impacts the effectiveness of VQG systems in applications like education and human-computer interaction, where appropriately challenging questions can stimulate learning interest and improve interaction experiences. However, accurately defining and controlling question difficulty is a challenging task due to its multidimensional and subjective nature. In this paper, we propose a new definition of the difficulty of questions, i.e., being positively correlated with the number of reasoning steps required to answer a question. For our definition, we construct a corresponding dataset and propose a benchmark as a foundation for future research. Our benchmark is designed to progressively increase the reasoning steps involved in generating questions. Specifically, we first extract the relationships among objects in the image to form a reasoning chain, then gradually increase the difficulty by rewriting the generated question to include more reasoning sub-chains. Experimental results on our constructed dataset show that our benchmark significantly outperforms existing baselines in controlling the reasoning chains of generated questions, producing questions with varying difficulty levels.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- PgM: Partitioner Guided Modal Learning FrameworkGuimin Hu, Yi Xin, Lijie Hu, Zhihong Zhu et al.ACM MM 2025 · 1 citation
- Diagram-Driven Course Questions GenerationXinyu Zhang, Lingling Zhang, Yanrui Wu, Muye Huang et al.EMNLP 2025
Builds on6
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- IMAGPose: A Unified Conditional Framework for Pose-Guided Person GenerationFei Shen, Jinhui TangNeurIPS 2024 · 172 citations
- Hierarchical Graph Network for Multi-hop Question AnsweringYuwei Fang, Siqi Sun, Zhe Gan, Rohit Pillai et al.EMNLP 2020 · 157 citations
- NLX-GPT: A Model for Natural Language Explanations in Vision and Vision-Language TasksFawaz Sammani, Tanmoy Mukherjee, Nikos DeligiannisCVPR 2022 · 46 citations
- Multiple Objects-Aware Visual Question GenerationJiayuan Xie, Yi Cai, Qingbao Huang, Tao WangACM MM 2021 · 23 citations
Related papers
- Guiding the Growth: Difficulty-Controllable Question Generation through Step-by-Step RewritingYi Cheng, Siyao Li, Bang Liu, Ruihui Zhao et al.ACL 2021
- Inferential Visual Question GenerationChao Bi, Shuhui Wang, Zhe Xue, Shengbo Chen et al.ACM MM 2022 · 8 citations
- On the General Value of Evidence, and Bilingual Scene-Text Visual Question AnsweringXinyu Wang, Yuliang Liu, Chunhua Shen, Chun Chet Ng et al.CVPR 2020
- Multi-VQG: Generating Engaging Questions for Multiple ImagesMin-Hsuan Yeh, Vincent Chen, Ting-Hao 'Kenneth' Huang, Lun-Wei KuEMNLP 2022 · 3 citations
- JRDB-Reasoning: A Difficulty-Graded Benchmark for Visual Reasoning in RoboticsSimindokht Jahangard, Mehrzad Mohammadi, Yi Shen, Zhixi Cai et al.AAAI 2026 · 2 citations
