Multiple Objects-Aware Visual Question Generation
Jiayuan Xie, Yi Cai, Qingbao Huang, Tao Wang
Abstract
Visual question generation task aims to generate meaningful questions about an image according to a target answer. Existing studies mainly focus on merely one object related to the target answer in an image to generate a question. However, a target answer is often related to multiple key objects in an image, which focuses on only one object may mislead its model to generate questions that are only related to partial fragments of the answer. To address this problem, we propose a multi-objects aware generation model to capture all key objects related to an answer and generate the corresponding question. We first introduce a co-attention network to capture the relationship between each object in an image and the answer, and then extract the key objects that are related to the answer. Then, a graph network is introduced to capture the relationships between the key objects and other objects in the image that are not related to the answer, which helps generate questions that involve more visual content. Finally, the learned information from the graph network is fed into a standard decoder module to produce questions. Extensive experiments on the VQA v2.0 dataset show that the proposed model outperforms the state-of-the-art models.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 5c22a8c5-042e-4455-ae2e-87e0ab477bc2Cited by top-tier papers3
- Surgment: Segmentation-enabled Semantic Search and Creation of Visual Question and Feedback to Support Video-Based Surgery LearningJingying Wang, Haoran Tang, Taylor Kantor, Tandis Soltani et al.CHI 2024 · 9 citations
- ConVQG: Contrastive Visual Question Generation with Multimodal GuidanceLi Mi, Syrielle Montariol, Javiera Castillo Navarro, Xianjie Dai et al.AAAI 2024 · 8 citations
- Explicitly Guided Difficulty-Controllable Visual Question GenerationJiayuan Xie, Mengqiu Cheng, Xinting Zhang, Yi Cai et al.AAAI 2025 · 2 citations
Related papers
- Re-Attention for Visual Question AnsweringWenya Guo, Ying Zhang, Xiaoping Wu, Jufeng Yang et al.AAAI 2020 · 90 citations
- Multi-Question Learning for Visual Question AnsweringChenyi Lei, Lei Wu, Dong Liu, Zhao Li et al.AAAI 2020 · 9 citations
- Deconfounded Visual Question Generation with Causal InferenceJiali Chen, Zhenjun Guo, Jiayuan Xie, Yi Cai et al.ACM MM 2023 · 8 citations
- Relation-Aware Graph Attention Network for Visual Question AnsweringLinjie Li, Zhe Gan, Yu Cheng, Jingjing LiuICCV 2019 · 391 citations
- Progressive Graph Attention Network for Video Question AnsweringLiang Peng, Shuangji Yang, Yi Bin, Guoqing WangACM MM 2021 · 47 citations
