Found a Reason for me? Weakly-supervised Grounded Visual Question Answering using Capsules
Aisha Urooj Khan, Hilde Kuehne, Kevin Duarte, Chuang Gan, Niels da Vitoria Lobo, Mubarak Shah
摘要
The problem of grounding VQA tasks has seen an increased attention in the research community recently, with most attempts usually focusing on solving this task by using pretrained object detectors. However, pre-trained object detectors require bounding box annotations for detecting relevant objects in the vocabulary, which may not always be feasible for real-life large-scale applications. In this paper, we focus on a more relaxed setting: the grounding of relevant visual entities in a weakly supervised manner by training on the VQA task alone. To address this problem, we propose a visual capsule module with a query-based selection mechanism of capsule features, that allows the model to focus on relevant regions based on the textual cues about visual information in the question. We show that integrating the proposed capsule module in existing VQA systems significantly improves their performance on the weakly supervised grounding task. Overall, we demonstrate the effectiveness of our approach on two state-of-the-art VQA systems, stacked NMN and MAC, on the CLEVR-Answers benchmark, our new evaluation set based on CLEVR scenes with groundtruth bounding boxes for objects that are relevant for the correct answer, as well as on GQA, a real world VQA dataset with compositional questions. We show that the systems with the proposed capsule module consistently outperform the respective baseline systems in terms of answer grounding, while achieving comparable performance on VQA task. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Grounding Answers for Visual Questions Asked by Visually Impaired PeopleChongyan Chen, Samreen Anjum, Danna GurariCVPR 2022 · 被引用 48 次
- Can I Trust Your Answer? Visually Grounded Video Question AnsweringJunbin Xiao, Angela Yao, Yicong Li, Tat-Seng ChuaCVPR 2024 · 被引用 44 次
- VQA Therapy: Exploring Answer Differences by Visually Grounding AnswersChongyan Chen, Samreen Anjum, Danna GurariICCV 2023 · 被引用 20 次
- Sentence Attention Blocks for Answer GroundingSeyedalireza Khoshsirat, Chandra KambhamettuICCV 2023 · 被引用 8 次
- Weakly-Supervised Generation and Grounding of Visual Descriptions with Conditional Generative ModelsEffrosyni Mavroudi, René VidalCVPR 2022 · 被引用 5 次
它引用的顶会 Paper9
- CLEVRER: Collision Events for Video Representation and ReasoningKexin Yi, Chuang Gan, Yunzhu Li, Pushmeet Kohli 等ICLR 2020 · 被引用 584 次
- Language-Conditioned Graph Networks for Relational ReasoningRonghang Hu, Anna Rohrbach, Trevor Darrell, Kate SaenkoICCV 2019 · 被引用 183 次
- TVQA+: Spatio-Temporal Grounding for Video Question AnsweringJie Lei, Licheng Yu, Tamara L. Berg, Mohit BansalACL 2020 · 被引用 173 次
- CapsuleVOS: Semi-Supervised Video Object Segmentation Using Capsule RoutingKevin Duarte, Yogesh S. Rawat, Mubarak ShahICCV 2019 · 被引用 80 次
- Detecting and Diagnosing Adversarial Images with Class-Conditional Capsule ReconstructionsYao Qin, Nicholas Frosst, Sara Sabour, Colin Raffel 等ICLR 2020 · 被引用 76 次
相关 Paper
- Pseudo-Q: Generating Pseudo Language Queries for Visual GroundingHaojun Jiang, Yuanze Lin, Dongchen Han, Shiji Song 等CVPR 2022 · 被引用 60 次
- Relation-aware Instance Refinement for Weakly Supervised Visual GroundingYongfei Liu, Bo Wan, Lin Ma, Xuming HeCVPR 2021
- Multimodal Neural Graph Memory Networks for Visual Question AnsweringMahmoud KhademiACL 2020 · 被引用 35 次
- Confidence-aware Pseudo-label Learning for Weakly Supervised Visual GroundingYang Liu, Jiahua Zhang, Qingchao Chen, Yuxin PengICCV 2023 · 被引用 19 次
- Distributed Attention for Grounded Image CaptioningNenglun Chen, Xingjia Pan, Runnan Chen, Lei Yang 等ACM MM 2021 · 被引用 19 次
