Transform-Retrieve-Generate: Natural Language-Centric Outside-Knowledge Visual Question Answering
Feng Gao, Qing Ping, Govind Thattai, Aishwarya N. Reganti, Ying Nian Wu, Prem Natarajan
Abstract
Outside-knowledge visual question answering (OK-VQA) requires the agent to comprehend the image, make use of relevant knowledge from the entire web, and digest all the information to answer the question. Most previous works address the problem by first fusing the image and question in the multi-modal space, which is inflexible for further fusion with a vast amount of external knowledge. In this paper, we call for an alternative paradigm for the OK-VQA task, which transforms the image into plain text, so that we can enable knowledge passage retrieval, and generative question-answering in the natural language space. This paradigm takes advantage of the sheer volume of gigantic knowledge bases and the richness of pretrained language models. A Transform-Retrieve-Generate framework (TRiG) framework is proposed <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup> <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup> The code of this work will be made public., which can be plug-and-played with alternative image-to-text models and textual knowledge bases. Experimental results show that our TRiG framework outperforms all state-of-the-art supervised methods by at least 11.1 % absolute margin.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 940367d9-3181-45be-b449-4ec9fd99d68aCited by top-tier papers21
- ViperGPT: Visual Inference via Python Execution for ReasoningDídac Surís, Sachit Menon, Carl VondrickICCV 2023 · 732 citations
- Fine-grained Late-interaction Multi-modal Retrieval for Retrieval Augmented Visual Question AnsweringWeizhe Lin, Jinghong Chen, Jingbiao Mei, Alexandru Coca et al.NeurIPS 2023 · 108 citations
- PromptCap: Prompt-Guided Image Captioning for VQA with GPT-3Yushi Hu, Hang Hua, Zhengyuan Yang, Weijia Shi et al.ICCV 2023 · 91 citations
- Visual Chain-of-Thought Prompting for Knowledge-Based Visual ReasoningZhenfang Chen, Qinhong Zhou, Yikang Shen, Yining Hong et al.AAAI 2024 · 77 citations
- CoTDet: Affordance Knowledge Prompting for Task Driven Object DetectionJiajin Tang, Ge Zheng, Jingyi Yu, Sibei YangICCV 2023 · 47 citations
Builds on12
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Large-Scale Adversarial Training for Vision-and-Language Representation LearningZhe Gan, Yen-Chun Chen, Linjie Li, Chen Zhu et al.NeurIPS 2020 · 561 citations
- An Empirical Study of GPT-3 for Few-Shot Knowledge-Based VQAZhengyuan Yang, Zhe Gan, Jianfeng Wang, Xiaowei Hu et al.AAAI 2022 · 517 citations
- Distilling Knowledge from Reader to Retriever for Question AnsweringGautier Izacard, Edouard GraveICLR 2021 · 317 citations
- Graph-Based Reasoning over Heterogeneous External Knowledge for Commonsense Question AnsweringShangwen Lv, Daya Guo, Jingjing Xu, Duyu Tang et al.AAAI 2020 · 224 citations
Related papers
- Retrieval Augmented Visual Question Answering with Outside KnowledgeWeizhe Lin, Bill ByrneEMNLP 2022 · 49 citations
- Combo of Thinking and Observing for Outside-Knowledge VQAQingyi Si, Yuchen Mo, Zheng Lin, Huishan Ji et al.ACL 2023 · 7 citations
- Entity-Focused Dense Passage Retrieval for Outside-Knowledge Visual Question AnsweringJialin Wu, Raymond J. MooneyEMNLP 2022 · 9 citations
- Retrieval-Augmented Visual Question Answering via Built-in Autoregressive Search EnginesXinwei Long, Zhiyuan Ma, Ermo Hua, Kaiyan Zhang et al.AAAI 2025 · 18 citations
- Multi-Modal Answer Validation for Knowledge-Based VQAJialin Wu, Jiasen Lu, Ashish Sabharwal, Roozbeh MottaghiAAAI 2022 · 183 citations
