Transform-Retrieve-Generate: Natural Language-Centric Outside-Knowledge Visual Question Answering
Feng Gao, Qing Ping, Govind Thattai, Aishwarya N. Reganti, Ying Nian Wu, Prem Natarajan
摘要
Outside-knowledge visual question answering (OK-VQA) requires the agent to comprehend the image, make use of relevant knowledge from the entire web, and digest all the information to answer the question. Most previous works address the problem by first fusing the image and question in the multi-modal space, which is inflexible for further fusion with a vast amount of external knowledge. In this paper, we call for an alternative paradigm for the OK-VQA task, which transforms the image into plain text, so that we can enable knowledge passage retrieval, and generative question-answering in the natural language space. This paradigm takes advantage of the sheer volume of gigantic knowledge bases and the richness of pretrained language models. A Transform-Retrieve-Generate framework (TRiG) framework is proposed <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup> <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup> The code of this work will be made public., which can be plug-and-played with alternative image-to-text models and textual knowledge bases. Experimental results show that our TRiG framework outperforms all state-of-the-art supervised methods by at least 11.1 % absolute margin.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper21
- ViperGPT: Visual Inference via Python Execution for ReasoningDídac Surís, Sachit Menon, Carl VondrickICCV 2023 · 被引用 732 次
- Fine-grained Late-interaction Multi-modal Retrieval for Retrieval Augmented Visual Question AnsweringWeizhe Lin, Jinghong Chen, Jingbiao Mei, Alexandru Coca 等NeurIPS 2023 · 被引用 108 次
- PromptCap: Prompt-Guided Image Captioning for VQA with GPT-3Yushi Hu, Hang Hua, Zhengyuan Yang, Weijia Shi 等ICCV 2023 · 被引用 91 次
- Visual Chain-of-Thought Prompting for Knowledge-Based Visual ReasoningZhenfang Chen, Qinhong Zhou, Yikang Shen, Yining Hong 等AAAI 2024 · 被引用 77 次
- CoTDet: Affordance Knowledge Prompting for Task Driven Object DetectionJiajin Tang, Ge Zheng, Jingyi Yu, Sibei YangICCV 2023 · 被引用 47 次
它引用的顶会 Paper12
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Large-Scale Adversarial Training for Vision-and-Language Representation LearningZhe Gan, Yen-Chun Chen, Linjie Li, Chen Zhu 等NeurIPS 2020 · 被引用 561 次
- An Empirical Study of GPT-3 for Few-Shot Knowledge-Based VQAZhengyuan Yang, Zhe Gan, Jianfeng Wang, Xiaowei Hu 等AAAI 2022 · 被引用 517 次
- Distilling Knowledge from Reader to Retriever for Question AnsweringGautier Izacard, Edouard GraveICLR 2021 · 被引用 317 次
- Graph-Based Reasoning over Heterogeneous External Knowledge for Commonsense Question AnsweringShangwen Lv, Daya Guo, Jingjing Xu, Duyu Tang 等AAAI 2020 · 被引用 224 次
相关 Paper
- Retrieval Augmented Visual Question Answering with Outside KnowledgeWeizhe Lin, Bill ByrneEMNLP 2022 · 被引用 49 次
- Combo of Thinking and Observing for Outside-Knowledge VQAQingyi Si, Yuchen Mo, Zheng Lin, Huishan Ji 等ACL 2023 · 被引用 7 次
- Entity-Focused Dense Passage Retrieval for Outside-Knowledge Visual Question AnsweringJialin Wu, Raymond J. MooneyEMNLP 2022 · 被引用 9 次
- Retrieval-Augmented Visual Question Answering via Built-in Autoregressive Search EnginesXinwei Long, Zhiyuan Ma, Ermo Hua, Kaiyan Zhang 等AAAI 2025 · 被引用 18 次
- Multi-Modal Answer Validation for Knowledge-Based VQAJialin Wu, Jiasen Lu, Ashish Sabharwal, Roozbeh MottaghiAAAI 2022 · 被引用 183 次
