Multiple Objects-Aware Visual Question Generation
Jiayuan Xie, Yi Cai, Qingbao Huang, Tao Wang
摘要
Visual question generation task aims to generate meaningful questions about an image according to a target answer. Existing studies mainly focus on merely one object related to the target answer in an image to generate a question. However, a target answer is often related to multiple key objects in an image, which focuses on only one object may mislead its model to generate questions that are only related to partial fragments of the answer. To address this problem, we propose a multi-objects aware generation model to capture all key objects related to an answer and generate the corresponding question. We first introduce a co-attention network to capture the relationship between each object in an image and the answer, and then extract the key objects that are related to the answer. Then, a graph network is introduced to capture the relationships between the key objects and other objects in the image that are not related to the answer, which helps generate questions that involve more visual content. Finally, the learned information from the graph network is fed into a standard decoder module to produce questions. Extensive experiments on the VQA v2.0 dataset show that the proposed model outperforms the state-of-the-art models.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper3
- Surgment: Segmentation-enabled Semantic Search and Creation of Visual Question and Feedback to Support Video-Based Surgery LearningJingying Wang, Haoran Tang, Taylor Kantor, Tandis Soltani 等CHI 2024 · 被引用 9 次
- ConVQG: Contrastive Visual Question Generation with Multimodal GuidanceLi Mi, Syrielle Montariol, Javiera Castillo Navarro, Xianjie Dai 等AAAI 2024 · 被引用 8 次
- Explicitly Guided Difficulty-Controllable Visual Question GenerationJiayuan Xie, Mengqiu Cheng, Xinting Zhang, Yi Cai 等AAAI 2025 · 被引用 2 次
相关 Paper
- Re-Attention for Visual Question AnsweringWenya Guo, Ying Zhang, Xiaoping Wu, Jufeng Yang 等AAAI 2020 · 被引用 90 次
- Multi-Question Learning for Visual Question AnsweringChenyi Lei, Lei Wu, Dong Liu, Zhao Li 等AAAI 2020 · 被引用 9 次
- Deconfounded Visual Question Generation with Causal InferenceJiali Chen, Zhenjun Guo, Jiayuan Xie, Yi Cai 等ACM MM 2023 · 被引用 8 次
- Relation-Aware Graph Attention Network for Visual Question AnsweringLinjie Li, Zhe Gan, Yu Cheng, Jingjing LiuICCV 2019 · 被引用 391 次
- Progressive Graph Attention Network for Video Question AnsweringLiang Peng, Shuangji Yang, Yi Bin, Guoqing WangACM MM 2021 · 被引用 47 次
