Prompting Large Language Models with Answer Heuristics for Knowledge-Based Visual Question Answering
Zhenwei Shao, Zhou Yu, Meng Wang, Jun Yu
摘要
Knowledge-based visual question answering (VQA) requires external knowledge beyond the image to answer the question. Early studies retrieve required knowledge from explicit knowledge bases (KBs), which often introduces irrelevant information to the question, hence restricting the performance of their models. Recent works have sought to use a large language model (i.e., GPT-3 [3]) as an implicit knowledge engine to acquire the necessary knowledge for answering. Despite the encouraging results achieved by these methods, we argue that they have not fully activated the capacity of GPT-3 as the provided input information is insufficient. In this paper, we present Prophet-a conceptually simple framework designed to prompt GPT-3 with answer heuristics for knowledge-based VQA. Specifically, we first train a vanilla VQA model on a specific knowledgebased VQA dataset without external knowledge. After that, we extract two types of complementary answer heuristics from the model: answer candidates and answer-aware examples. Finally, the two types of answer heuristics are encoded into the prompts to enable GPT-3 to better comprehend the task thus enhancing its capacity. Prophet significantly outperforms all existing state-of-the-art methods on two challenging knowledge-based VQA datasets, OK-VQA and A-OKVQA, delivering 61.1% and 55.7% accuracies on their testing sets, respectively. Q: what fills the balloons? testing sample training samples answer-aware examples Candidates: • candle(0.99) • birthday(0.02) • fire(0.01) ...
helium Please answer the question according to the context and answer candidates . Each answer candidate is associated with a confidence score within a bracket. The true answer may not be included in the candidates. Candidates: • air(0.28) • helium(0.07) • wine(0.03) Candidates: • air(0.69) • helium(0.62) • oxygen(0.04) ... Frozen GPT-3 Model Context: Inflated kites in various shapes float in the air. Question: What chemical makes cats fly? Candidates: air (0.69), helium (0.62), oxygen (0.04) Answer: helium Context: The man is smiling at a birthday cake. Question: What is he about to blow out? Candidates: candle (0.99), birthday (0.02), fire (0.01) Answer: candle Context: a group of children stand around a cake. Question: What fills the balloons? Candidates: air (0.28), helium (0.07), wine (0.03
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper57
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong 等NeurIPS 2023 · 被引用 4,013 次
- Fine-grained Late-interaction Multi-modal Retrieval for Retrieval Augmented Visual Question AnsweringWeizhe Lin, Jinghong Chen, Jingbiao Mei, Alexandru Coca 等NeurIPS 2023 · 被引用 108 次
- Large Language Models are Visual Reasoning CoordinatorsLiangyu Chen, Bo Li, Sheng Shen, Jingkang Yang 等NeurIPS 2023 · 被引用 108 次
- KAM-CoT: Knowledge Augmented Multimodal Chain-of-Thoughts ReasoningDebjyoti Mondal, Suraj Modi, Subhadarshi Panda, Rituraj Singh 等AAAI 2024 · 被引用 96 次
- Visual Chain-of-Thought Prompting for Knowledge-Based Visual ReasoningZhenfang Chen, Qinhong Zhou, Yikang Shen, Yining Hong 等AAAI 2024 · 被引用 77 次
它引用的顶会 Paper22
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech 等NeurIPS 2022 · 被引用 6,707 次
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 被引用 6,549 次
- OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning FrameworkPeng Wang, An Yang, Rui Men, Junyang Lin 等ICML 2022 · 被引用 1,058 次
相关 Paper
- TOA: Task-oriented Active VQAXiaoying Xing, Mingfu Liang, Ying WuNeurIPS 2023 · 被引用 20 次
- PromptCap: Prompt-Guided Image Captioning for VQA with GPT-3Yushi Hu, Hang Hua, Zhengyuan Yang, Weijia Shi 等ICCV 2023 · 被引用 91 次
- An Empirical Study of GPT-3 for Few-Shot Knowledge-Based VQAZhengyuan Yang, Zhe Gan, Jianfeng Wang, Xiaowei Hu 等AAAI 2022 · 被引用 517 次
- Transform-Retrieve-Generate: Natural Language-Centric Outside-Knowledge Visual Question AnsweringFeng Gao, Qing Ping, Govind Thattai, Aishwarya N. Reganti 等CVPR 2022 · 被引用 85 次
- Breaking the Barrier Between Pre-training and Fine-tuning: A Hybrid Prompting Model for Knowledge-Based VQAZhongfan Sun, Yongli Hu, Qingqing Gao, Huajie Jiang 等ACM MM 2023 · 被引用 6 次
