VQA With No Questions-Answers Training
Ben Zion Vatashsky, Shimon Ullman
Abstract
Methods for teaching machines to answer visual questions have made significant progress in recent years, but current methods still lack important human capabilities, including integrating new visual classes and concepts in a modular manner, providing explanations for the answers and handling new domains without explicit examples. We propose a novel method that consists of two main parts: generating a question graph representation, and an answering procedure, guided by the abstract structure of the question graph to invoke an extendable set of visual estimators. Training is performed for the language part and the visual part on their own, but unlike existing schemes, the method does not require any training using images with associated questions and answers. This approach is able to handle novel domains (extended question types and new object classes, properties and relations) as long as corresponding visual estimators are available. In addition, it can provide explanations to its answers and suggest alternatives when questions are not grounded in the image. We demonstrate that this approach achieves both high performance and domain extensibility without any questions-answers training.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Just Ask: Learning to Answer Questions from Millions of Narrated VideosAntoine Yang, Antoine Miech, Josef Sivic, Ivan Laptev et al.ICCV 2021 · 345 citations
- Measuring Compositional Consistency for Video Question AnsweringMona Gandhi, Mustafa Omer Gul, Eva Prakash, Madeleine Grunde-McLaughlin et al.CVPR 2022 · 11 citations
- Domain-Robust VQA With Diverse Datasets and Methods but No Target LabelsMingda Zhang, Tristan Maidment, Ahmad Diab, Adriana Kovashka et al.CVPR 2021
- AGQA: A Benchmark for Compositional Spatio-Temporal ReasoningMadeleine Grunde-McLaughlin, Ranjay Krishna, Maneesh AgrawalaCVPR 2021
Builds on4
- Unified Vision-Language Pre-Training for Image Captioning and VQALuowei Zhou, Hamid Palangi, Lei Zhang, Houdong Hu et al.AAAI 2020 · 1,047 citations
- Relation-Aware Graph Attention Network for Visual Question AnsweringLinjie Li, Zhe Gan, Yu Cheng, Jingjing LiuICCV 2019 · 391 citations
- Taking a HINT: Leveraging Explanations to Make Vision and Language Models More GroundedRamprasaath Ramasamy Selvaraju, Stefan Lee, Yilin Shen, Hongxia Jin et al.ICCV 2019 · 288 citations
- Compact Trilinear Interaction for Visual Question AnsweringTuong Do, Huy Tran, Thanh-Toan Do, Erman Tjiputra et al.ICCV 2019 · 63 citations
Related papers
- Counterfactual Vision and Language LearningEhsan Abbasnejad, Damien Teney, Amin Parvaneh, Javen Shi et al.CVPR 2020
- Why Does a Visual Question Have Different Answers?Nilavra Bhattacharya, Qing Li, Danna GurariICCV 2019 · 78 citations
- SGEITL: Scene Graph Enhanced Image-Text Learning for Visual Commonsense ReasoningZhecan Wang, Haoxuan You, Liunian Harold Li, Alireza Zareian et al.AAAI 2022 · 40 citations
- Separating Skills and Concepts for Novel Visual Question AnsweringSpencer Whitehead, Hui Wu, Heng Ji, Rogério Feris et al.CVPR 2021
- How Can Objects Help Video-Language Understanding?Zitian Tang, Shijie Wang, Junho Cho, Jaewook Yoo et al.ICCV 2025 · 8 citations
